Courseiva

SAA-C03 (SAA-C03) — Questions 601–675

935 questions total · 13pages · All types, answers revealed

Page 8

Page 9 of 13

Page 10
601
Multi-Selecthard

A serverless checkout API runs on AWS Lambda behind API Gateway. Traffic spikes are predictable every weekday at 09:00 UTC, and p95 latency jumps for the first few minutes after each deployment because execution environments are cold. The team wants to reduce this startup impact without changing the API contract. Which changes should they make? Select three.

Select 3 answers
A.Configure provisioned concurrency on the production Lambda alias during the busy windows.
B.Initialize SDK clients and other reusable objects outside the handler so they are created once per execution environment.
C.Reduce the deployment package size and remove unnecessary layers to shorten function initialization.
D.Replace provisioned concurrency with reserved concurrency because reserved concurrency keeps instances warm.
E.Increase the function timeout so the first request has more time to warm up.
AnswersA, B, C

Correct. Provisioned concurrency keeps a pool of pre-initialized execution environments ready to handle invocations, which directly reduces cold-start latency. Using an alias allows the team to manage production traffic separately from development or canary versions and to schedule capacity for the predictable weekday peak.

Why this answer

Provisioned concurrency initializes a specified number of execution environments in advance, so when traffic spikes at 09:00 UTC, the Lambda function is already warm and can serve requests without cold start latency. This directly addresses the p95 latency jump after deployment without altering the API contract.

Exam trap

The trap here is confusing reserved concurrency (which only limits concurrency) with provisioned concurrency (which pre-warms instances), leading candidates to incorrectly select reserved concurrency as a solution for cold starts.

602
MCQmedium

A production internal reporting portal runs continuously on EC2 with predictable usage for the next three years. The team wants a discount while retaining some instance-family flexibility. What should they buy? The design must avoid adding custom operational scripts.

A.Spot Instances only
B.Dedicated Instances
C.Compute Savings Plan
D.S3 Intelligent-Tiering
AnswerC

A Compute Savings Plan offers significant discounted rates in exchange for a commitment to a consistent amount of compute usage (e.g., $/hour) over a 1- or 3-year term. It automatically covers any EC2 instance family, size, or Region, as well as Fargate and Lambda, providing flexibility if the reporting portal's instance mix changes over time. This makes it the ideal solution for a continuously running production workload because it delivers up to a 66% discount versus On-Demand without locking you into specific instance specs.

Why this answer

Compute Savings Plans offer the lowest prices on EC2 instance usage (up to 66% off On-Demand) in exchange for a 1- or 3-year commitment, while allowing flexibility across instance families, sizes, OS, and regions. This matches the requirement for a discount on predictable, continuous usage without locking into a specific instance type, and requires no custom scripts.

Exam trap

The trap here is that candidates confuse Savings Plans with Reserved Instances, assuming they require instance-family lock-in, or they incorrectly apply storage services (S3 Intelligent-Tiering) to compute cost optimization.

How to eliminate wrong answers

Option A is wrong because Spot Instances are not suitable for a continuously running production portal; they can be interrupted with a 2-minute warning, making them unreliable for steady-state workloads. Option B is wrong because Dedicated Instances provide physical isolation at a higher cost and do not offer a discount mechanism; they are for compliance or licensing needs, not cost savings. Option D is wrong because S3 Intelligent-Tiering is an object storage class for data with changing access patterns, not applicable to EC2 compute instances.

603
MCQmedium

A website serves mostly cacheable images, CSS, and JavaScript from an ALB. Users in Europe and Asia report slower page loads, and the ALB receives far more requests than expected. The team also wants text assets compressed automatically. Which change is the best first step?

A.Increase the ALB size and add more target instances behind it.
B.Use Route 53 latency-based routing to send users to the nearest ALB.
C.Place Amazon CloudFront in front of the ALB and enable compression and caching.
D.Replace the ALB with an NLB to reduce latency for web requests.
AnswerC

CloudFront is the right choice because it caches static content at edge locations close to users, reducing latency and lowering the number of requests that reach the ALB. It also supports compression for text-based assets such as CSS, JavaScript, and HTML. This improves both performance and origin offload without changing the application logic.

Why this answer

CloudFront is the correct first step because it acts as a CDN that caches cacheable content (images, CSS, JS) at edge locations close to users in Europe and Asia, reducing load on the ALB and improving page load times. It also supports automatic compression of text assets (e.g., via gzip or Brotli) without requiring backend changes, directly addressing the team's requirement for compressed text assets. By offloading requests from the ALB, CloudFront reduces the number of requests hitting the origin, solving the 'far more requests than expected' issue.

Exam trap

The trap here is that candidates often think scaling the ALB (Option A) or using latency-based routing (Option B) will solve performance issues, but they overlook that caching and compression at the edge (CloudFront) directly address both latency and request volume without requiring backend changes.

Why the other options are wrong

A

The issue is not ALB capacity or backend scaling; it's about excessive requests and latency due to lack of caching and compression. Increasing ALB size and instances doesn't reduce request volume or compress assets.

B

Route 53 latency-based routing directs users to the nearest ALB, but the problem is that the ALB receives far more requests than expected due to cacheable content not being cached. Latency routing alone does not reduce ALB load or compress text assets automatically.

D

An NLB does not support caching, compression, or HTTP-level features like image/CSS/JS optimization; it operates at Layer 4 and cannot reduce request volume or compress text assets.

When would these options actually be correct?

A

When the website experiences high CPU/memory utilization on the ALB and backend instances, and the primary bottleneck is compute capacity rather than request volume or latency, scaling out the ALB and targets is appropriate.

B

In a scenario where you have multiple ALBs deployed in different AWS regions and users experience high latency because they are being routed to a distant region, Route 53 latency-based routing would be the correct answer to direct each user to the closest ALB, reducing latency.

D

If the question required handling millions of UDP or TCP connections with ultra-low latency and no need for HTTP features (e.g., a real-time gaming server or IoT device traffic), replacing an ALB with an NLB would be appropriate.

Why candidates pick the wrong answer

A

Candidates often default to scaling solutions when they see performance issues, without diagnosing the root cause (lack of caching/compression).

B

Candidates may think that reducing network latency by routing users to the nearest ALB will solve the slow page loads, but they overlook that the core issue is excessive requests hitting the ALB due to lack of caching and compression, which latency routing does not address.

D

Candidates may think NLB reduces latency because it is faster at the transport layer, but they overlook that the bottleneck here is cacheable content and compression, which require Layer 7 features.

604
MCQeasy

A system processes events from Amazon SQS and sometimes sees duplicate messages due to retries. The business requirement is that each payment must be charged at most once. What design choice best addresses this resiliency requirement?

A.Assume duplicates never occur because the consumer deletes messages immediately after receiving them.
B.Implement idempotent processing using a deduplication key (for example, paymentId) and record completed charges so duplicates are safely ignored.
C.Increase the SQS visibility timeout until duplicates never happen.
D.Use SNS topics instead of SQS so retries are disabled by default.
AnswerB

Idempotency ensures at-most-once side effects even when duplicates are delivered. Persist a record keyed by paymentId (e.g., a unique constraint/conditional write). If the record indicates the payment was already charged, skip the charge for any subsequent duplicate message.

Why this answer

Implementing idempotent processing with a deduplication key (e.g., paymentId) ensures that even if duplicate messages arrive from SQS (due to retries or at-least-once delivery), the consumer can check a record of completed charges and safely ignore duplicates. This satisfies the business requirement of charging each payment at most once without relying on SQS’s best-effort deduplication or message ordering.

Exam trap

The trap here is that candidates assume SQS guarantees exactly-once delivery or that increasing visibility timeouts can prevent duplicates, but SQS is designed for at-least-once delivery, and the only reliable way to handle duplicates is to make the consumer idempotent.

How to eliminate wrong answers

Option A is wrong because SQS provides at-least-once delivery, and deleting a message immediately after receiving it does not prevent duplicates that may arrive before the delete is processed or due to visibility timeout expiration; assuming duplicates never occur violates the fundamental reliability guarantee of SQS. Option C is wrong because increasing the visibility timeout cannot eliminate duplicates; it only delays the redelivery of unacknowledged messages, and duplicates can still occur due to network retries, consumer crashes, or SQS’s internal replication. Option D is wrong because SNS topics do not disable retries by default; SNS uses at-least-once delivery and can retry HTTP/S endpoints, and switching to SNS does not solve the duplicate problem—it may even introduce additional delivery attempts without built-in deduplication.

605
MCQmedium

A latency-sensitive API is implemented with AWS Lambda. The team enabled provisioned concurrency to avoid cold starts, setting provisioned concurrency to 50 because marketing campaigns occasionally cause spikes. However, during most weekdays the API receives little traffic (near zero), and the team is seeing high monthly Lambda costs from idle provisioned capacity. What is the best cost-optimized strategy that still meets the requirement of fast initial responses during traffic spikes?

A.Increase provisioned concurrency to 100 so that cold starts never occur, regardless of traffic patterns.
B.Use Application Auto Scaling scheduled actions to increase provisioned concurrency on the Lambda alias before campaign windows and reduce it to a minimal baseline afterward.
C.Turn provisioned concurrency off permanently and rely on retries at the client side to mask cold starts.
D.Replace Lambda with a single always-on EC2 instance sized for peak demand to eliminate cold starts.
AnswerB

Application Auto Scaling scheduled actions let you define time-based policies on a Lambda alias, so provisioned concurrency is raised to a high value just before the campaign starts and lowered to a minimal baseline once it ends. During the campaign, requests are served by pre-initialized execution environments, eliminating the cold-start latency that would otherwise be noticeable. Because provisioned concurrency is billed while allocated even when idle, the schedule ensures you only pay for warm capacity during the actual spike window, not during all other hours.

Why this answer

It uses Application Auto Scaling scheduled actions to dynamically adjust provisioned concurrency, scaling up to 50 before marketing campaigns and reducing to a minimal baseline (e.g., 1-5) during low-traffic weekdays. This eliminates idle capacity costs while ensuring fast initial responses during spikes, as provisioned concurrency keeps Lambda environments warm and ready to handle requests without cold starts.

Exam trap

The trap here is that candidates may assume provisioned concurrency must be set to a static high value to handle spikes, ignoring AWS's native Auto Scaling capabilities that can dynamically adjust capacity based on schedule or metrics, thus missing the cost-optimization aspect of the question.

How to eliminate wrong answers

Option A is wrong because increasing provisioned concurrency to 100 would double the idle capacity cost during low-traffic periods, exacerbating the cost issue without addressing the root problem of over-provisioning. Option C is wrong because turning off provisioned concurrency permanently would cause cold starts on every invocation during traffic spikes, violating the latency-sensitive requirement; client-side retries do not mask the initial latency of a cold start (typically 1-5 seconds for Lambda). Option D is wrong because replacing Lambda with a single always-on EC2 instance sized for peak demand would incur higher costs (24/7 compute) and eliminate the serverless benefits of automatic scaling and pay-per-use, while still risking performance degradation if the single instance is overwhelmed.

606
MCQeasy

A company hosts a web application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer (ALB). The ALB and the Auto Scaling group are currently deployed in only one Availability Zone (AZ). The business wants the application to keep running if that AZ has an outage. What is the best change?

A.Increase the desired capacity in the existing Availability Zone to handle all traffic during an outage.
B.Deploy the ALB and the Auto Scaling group across at least two Availability Zones so healthy targets remain.
C.Enable longer ALB health check intervals so failing instances are detected more slowly.
D.Switch from the ALB to an Internet Gateway so instances can fail over to the public internet.
AnswerB

To tolerate an AZ outage, both the load-balancing entry point (the ALB) and the compute capacity (the Auto Scaling instances) must be available in more than one AZ. With the ALB in multiple AZs and the Auto Scaling group using multiple subnets/AZs, requests can be routed to healthy targets in a remaining AZ while Auto Scaling replaces unhealthy instances.

Why this answer

Deploying the ALB and Auto Scaling group across at least two Availability Zones (AZs) ensures that if one AZ fails, the ALB can route traffic to healthy EC2 instances in the remaining AZ(s). This is the fundamental AWS best practice for high availability: an ALB is a regional service that requires targets in multiple AZs to survive an AZ outage, and the Auto Scaling group must also span those AZs to maintain capacity. Without multi-AZ deployment, a single AZ failure makes the entire application unavailable regardless of instance health checks.

Exam trap

The trap here is that candidates think increasing capacity or adjusting health check intervals can compensate for a single-AZ deployment, but AWS high availability fundamentally requires distributing resources across multiple isolated failure domains (AZs).

How to eliminate wrong answers

Option A is wrong because increasing the desired capacity in a single AZ does not protect against an AZ outage; all instances are in the same failure domain, so they all become unreachable simultaneously. Option C is wrong because enabling longer health check intervals would delay detection of failing instances, making the application less responsive to failures and increasing downtime, not improving availability. Option D is wrong because an Internet Gateway (IGW) is a VPC component that enables outbound internet access for instances, not a load balancer; it cannot perform health checks, distribute traffic, or fail over traffic between instances, and it does not replace the ALB's role in high availability.

607
Multi-Selectmedium

A containerized service on Amazon ECS connects to a database with a password that must never be stored in plaintext or hardcoded in the image. The application reads the password at startup and occasionally reconnects later, so it needs to retrieve the current secret when needed. Which three actions should the architect take? Select three.

Select 3 answers
A.Store the database password in AWS Secrets Manager.
B.Have the application retrieve the secret from Secrets Manager at runtime when it needs the password.
C.Grant the ECS task role least-privilege permission to read only that secret.
D.Store the password in a plain environment variable and update it manually during maintenance windows.
E.Use an IAM user access key inside the container so the database password can be embedded in code.
AnswersA, B, C

AWS Secrets Manager stores the credential encrypted with KMS and lets the ECS task retrieve it at runtime via the AWS SDK, so nothing is baked into the image or plaintext. This satisfies the requirement to fetch the current secret on each reconnect, and rotation keeps the retrieved value valid.

Why this answer

Option A is correct because AWS Secrets Manager is purpose-built to store sensitive values like database passwords encrypted at rest with KMS, so the password is never kept in plaintext or baked into the image. Option B is correct because the application must call the Secrets Manager API (e.g., GetSecretValue) at runtime each time it needs the credential, which also lets it pick up rotated values on later reconnects. Option C is correct because the ECS task role should be granted least-privilege IAM permissions scoped to just that secret's ARN (secretsmanager:GetSecretValue), so only the task can read it and nothing broader is exposed.

Option D is wrong because a plain environment variable stores the password in cleartext and manual updates are error-prone and insecure. Option E is wrong because embedding the password in code and using long-lived IAM user access keys violates the requirement to avoid hardcoding and least-privilege credential management.

Exam trap

The trap here is that candidates might think environment variables or IAM access keys are acceptable for secrets, but the exam requires using a dedicated secrets management service like Secrets Manager to avoid plaintext exposure and enable rotation.

Why the other options are wrong

D

Storing the password in a plain environment variable violates the requirement that the password must never be stored in plaintext. Manual updates during maintenance windows are not secure and do not provide automated rotation or retrieval at runtime.

E

Using an IAM user access key inside the container violates the principle of not storing secrets in the image or code, and access keys are long-lived credentials that increase security risk. The correct approach is to use IAM roles for tasks to obtain temporary credentials.

When would these options actually be correct?

D

In a non-production environment with no security compliance requirements, where the password is static and the application reads it from an environment variable set at container launch, and manual updates are acceptable for testing purposes.

E

In a scenario where an application needs to authenticate to an external API that requires long-lived access keys and the application is running on an EC2 instance without IAM roles support, embedding an IAM user access key in a secure configuration file (not in code) might be acceptable if encrypted and rotated regularly.

Why candidates pick the wrong answer

D

Candidates may think environment variables are a simple and acceptable way to pass secrets, overlooking the explicit requirement to avoid plaintext storage and the need for runtime retrieval without hardcoding.

E

Candidates may think that IAM access keys are a standard way to grant programmatic access, and they might not fully understand that ECS tasks can assume IAM roles, eliminating the need to embed keys.

608
MCQmedium

A telemetry pipeline uses an Application Load Balancer in one Region. Global users need lower network latency to the application without caching dynamic responses. What should be considered? The design must avoid adding custom operational scripts.

A.AWS Global Accelerator
B.S3 Cross-Region Replication
C.CloudFront only with long TTLs
D.AWS Backup cross-Region copy
AnswerA

Global Accelerator routes traffic over the AWS global network to improve performance for TCP/UDP applications without relying on caching.

Why this answer

AWS Global Accelerator uses the AWS global network to route traffic from global users to the Application Load Balancer, reducing latency and jitter by leveraging Anycast IP addresses and edge locations. It does not cache responses, making it ideal for dynamic content where low latency is required without custom operational scripts.

Exam trap

The trap here is that candidates often confuse CloudFront with Global Accelerator, assuming CloudFront can reduce latency for dynamic content without caching, but CloudFront inherently caches content at edge locations and requires custom origin headers or Lambda@Edge to bypass caching, which violates the 'no custom operational scripts' constraint.

How to eliminate wrong answers

Option B (S3 Cross-Region Replication) is wrong because it replicates objects across S3 buckets in different Regions, which does not reduce network latency for dynamic application traffic and is unrelated to ALB routing. Option C (CloudFront only with long TTLs) is wrong because CloudFront caches responses at edge locations, and long TTLs would serve stale dynamic content, contradicting the requirement to avoid caching dynamic responses. Option D (AWS Backup cross-Region copy) is wrong because it is a backup and disaster recovery service that copies backup data across Regions, not a solution for reducing network latency to an application endpoint.

609
MCQmedium

Account A has an IAM role named FinanceDataRole that is assumed by a principal in Account B. The role’s trust policy includes a condition requiring sts:ExternalId to equal "Fin-2026-Q2". A developer in Account B calls AssumeRole but receives an error: AccessDenied: ExternalId mismatch. The security team requires that you do not remove the ExternalId condition. What is the correct remediation?

A.Add kms:Decrypt to the developer’s IAM policy so KMS can validate the ExternalId during AssumeRole.
B.Update the AssumeRole call in Account B to include sts:ExternalId="Fin-2026-Q2" exactly as required.
C.Increase the role’s MaxSessionDuration to reduce authentication failures.
D.Remove the ExternalId condition from the trust policy to allow all AssumeRole requests.
AnswerB

To satisfy the role's trust policy, the AssumeRole request from Account B must explicitly include the parameter sts:ExternalId with the exact value "Fin-2026-Q2". STS evaluates this parameter against the StringEquals condition on sts:ExternalId in the trust policy; if it matches, the policy condition passes and STS issues temporary credentials. Supplying any other value or omitting the parameter will cause the AssumeRole call to be denied.

Why this answer

The error 'AccessDenied: ExternalId mismatch' occurs because the AssumeRole API call from Account B does not include the required sts:ExternalId parameter. The trust policy on the FinanceDataRole explicitly requires this parameter to match 'Fin-2026-Q2' as a security measure to prevent the confused deputy problem. Option B is correct because the developer must pass the exact ExternalId value in the AssumeRole request to satisfy the condition and successfully assume the role.

Exam trap

The trap here is that candidates may think the ExternalId is automatically passed or that the error is due to permission issues (like KMS or session duration), when in fact the developer must explicitly include the correct ExternalId in the AssumeRole API call.

How to eliminate wrong answers

Option A is wrong because KMS is not involved in validating ExternalId during AssumeRole; the ExternalId check is performed by the AWS STS service based on the role's trust policy, not by KMS. Option C is wrong because MaxSessionDuration controls the maximum session length for an assumed role, not authentication failures related to ExternalId mismatches. Option D is wrong because the security team explicitly requires that the ExternalId condition not be removed, and removing it would weaken security by eliminating the confused deputy protection.

610
MCQeasy

A company has a primary application in us-east-1 and a standby environment in us-west-2. Users should go to the primary site while it is healthy and automatically switch to the standby site if the primary fails. Which Route 53 routing policy should they use?

A.Weighted routing
B.Failover routing with health checks
C.Geolocation routing
D.Latency-based routing
AnswerB

Route 53 failover routing is designed for active-passive resilience. You define a primary record and a secondary record, then attach health checks so DNS answers shift to the standby when the primary is unhealthy. This provides a simple disaster recovery pattern for user-facing endpoints without requiring application-level traffic management.

Why this answer

Failover routing with health checks is the correct choice because it allows you to configure an active-passive failover pattern where Route 53 directs traffic to the primary resource (us-east-1) as long as it passes a health check. If the health check fails, Route 53 automatically routes traffic to the secondary resource (us-west-2), ensuring high availability without manual intervention.

Exam trap

The trap here is that candidates often confuse failover routing with latency-based routing, thinking that latency routing will automatically switch to a healthy region, but latency routing only optimizes for speed and does not consider health status unless combined with health checks, which is not its primary purpose.

Why the other options are wrong

A

Weighted routing distributes traffic across multiple resources based on assigned weights, but it does not automatically fail over to a standby site when the primary fails; it requires manual weight adjustment or health checks to redirect traffic.

C

Geolocation routing directs traffic based on the user's geographic location, not on the health of the primary site. It cannot automatically fail over to a standby region when the primary fails.

D

Latency-based routing directs users to the region with the lowest latency, not to a primary site with automatic failover to a standby site. It does not support health checks to detect primary failure.

When would these options actually be correct?

A

A company wants to send 10% of traffic to a new application version for testing while sending 90% to the stable version, using Route 53 to distribute traffic based on weights.

C

A company needs to route users to different endpoints based on their country or continent, e.g., users in Europe go to eu-west-1, users in North America go to us-east-1, and they want to serve localized content or comply with data residency requirements.

D

A company wants to route users to the AWS region that provides the lowest latency for each user, without a fixed primary/standby setup. For example, a global application serving users worldwide should use latency-based routing to minimize response times.

Why candidates pick the wrong answer

A

Candidates may think weighted routing can be used to send all traffic to the primary by setting its weight to 100 and the standby to 0, but this does not provide automatic failover without health checks.

C

Candidates may confuse geolocation with geographic failover, thinking that routing based on location can also handle failover, or they may assume that 'geolocation' implies awareness of site health.

D

Candidates may think latency-based routing can automatically redirect users when the primary fails, confusing performance optimization with disaster recovery failover.

611
MCQhard

Based on the exhibit, an application runs on Amazon Aurora MySQL. The writer instance is frequently near 85% CPU while the reader instance is under 20% CPU. Application traces show that most of the database traffic is read-only SELECT queries, but the code currently sends all queries to the writer endpoint. What should the solutions architect recommend to improve performance with the smallest functional change?

A.Increase the writer instance size and keep all traffic on the writer endpoint.
B.Point read-only database traffic to the Aurora reader endpoint and keep writes on the writer endpoint.
C.Convert the cluster to a Multi-AZ RDS PostgreSQL deployment to get automatic failover and better read performance.
D.Enable cross-Region read replicas so SELECT queries are routed to a remote Region for improved performance.
AnswerB

Routing read-only SELECTs to the Aurora reader endpoint offloads the reader instance already sitting under 20% CPU, while writes stay on the writer endpoint. This satisfies the smallest-functional-change constraint: only the connection endpoint for read traffic changes, requiring no schema, engine or application-logic redesign.

Why this answer

The Aurora reader endpoint distributes read-only traffic across all available reader instances, offloading the writer instance and reducing its CPU utilization. Since the application traces show most traffic is read-only SELECT queries, this change requires only modifying the connection string for reads while keeping writes on the writer endpoint, making it the smallest functional change.

Exam trap

The trap here is that candidates may think increasing instance size (Option A) is the simplest fix, but they overlook the fact that Aurora's architecture is designed to offload reads to reader instances, which is a more cost-effective and scalable solution with minimal code change.

How to eliminate wrong answers

Option A is wrong because increasing the writer instance size does not address the root cause—the writer is overloaded with read traffic that could be handled by readers—and it incurs higher cost without leveraging Aurora's built-in read scaling. Option C is wrong because converting to RDS PostgreSQL Multi-AZ does not provide the same read scaling as Aurora readers; Multi-AZ only provides a standby for failover, not active read offloading, and it requires a full migration. Option D is wrong because cross-Region read replicas introduce significant latency for read queries and are intended for disaster recovery or global read scaling, not for reducing CPU on the local writer instance.

612
MCQmedium

A public API for a B2B file exchange site is deployed on API Gateway. Clients must authenticate with standards-based tokens issued by an external OpenID Connect provider. Which authorization mechanism should be used?

A.API keys only
B.IAM authorization for all internet users
C.JWT authorizer configured for the OpenID Connect issuer
D.A VPC endpoint policy
AnswerC

A JWT authorizer in API Gateway validates the RS256 signature, issuer, and audience of a JSON Web Token issued by a trusted OpenID Connect provider, such as Auth0, Okta, or Amazon Cognito User Pools. This authorizer leverages the provider's JWKS endpoint to verify tokens without requiring a custom Lambda authorizer or backend authentication logic, significantly reducing operational overhead. Additionally, it can map claims from the token into the request context, enabling fine-grained routing or scope-based restrictions. This directly meets the need to authenticate public API users who already hold OIDC-issued access tokens.

Why this answer

API Gateway's JWT authorizer natively validates JSON Web Tokens issued by an external OpenID Connect (OIDC) provider. It verifies the token's signature, expiry, and issuer against the OIDC provider's JWKS endpoint without requiring custom Lambda code, making it the simplest and most secure choice for standards-based token authentication.

Exam trap

The trap here is that candidates often confuse API keys (which are for rate limiting and client identification) with authentication, or assume IAM authorization can be used for external users, but IAM requires AWS credentials and is not designed for third-party OIDC tokens.

How to eliminate wrong answers

Option A is wrong because API keys only provide client identification, not authentication; they do not validate the identity of the caller or support OIDC tokens. Option B is wrong because IAM authorization is designed for AWS principals (e.g., IAM users/roles) and requires AWS Signature V4 signing, which is not suitable for internet clients using external OIDC tokens. Option D is wrong because a VPC endpoint policy controls access to API Gateway via VPC endpoints, not authentication or token validation for public internet clients.

613
Multi-Selecthard

A product catalog system uses a relational database for orders and a simple key-value profile store for shopping carts. Traffic is unpredictable, and the company wants to avoid paying for large idle database instances. Which two choices are best? Select two.

Select 2 answers
A.Use Aurora Serverless v2 for the relational order system.
B.Use DynamoDB on-demand capacity for the shopping-cart profile store.
C.Keep both workloads on large provisioned RDS instances and add read replicas for the cart store.
D.Use DynamoDB provisioned capacity with a fixed minimum despite the unpredictable traffic.
E.Replace the relational order system with a wide-column table to reduce SQL licensing.
AnswersA, B

Aurora Serverless v2 scales capacity automatically in fine-grained increments based on actual load, so the relational order system pays only for consumed capacity during unpredictable traffic. This directly satisfies the requirement to avoid paying for large idle database instances without application rewrites.

Why this answer

A is correct because Aurora Serverless v2 automatically scales database capacity in fine-grained ACUs up and down based on actual load, so the relational order system only consumes (and bills for) the capacity it needs during unpredictable traffic, avoiding large idle provisioned instances. B is correct because DynamoDB on-demand capacity mode charges per request and instantly accommodates unpredictable traffic spikes without provisioning or managing capacity, which fits the key-value shopping-cart profile store. C is wrong because large provisioned RDS instances plus read replicas still pay for idle capacity and read replicas do not address write-side traffic variability.

D is wrong because DynamoDB provisioned capacity with a fixed minimum requires capacity planning and either throttles or wastes money under unpredictable traffic. E is wrong because replacing the relational order system with a wide-column store does not reduce SQL licensing (Aurora is not licensed per-core like commercial engines) and abandons the relational model the orders workload requires.

Exam trap

The trap here is that candidates may think provisioned capacity with a minimum is acceptable for unpredictable traffic, but the question explicitly requires avoiding paying for idle capacity, so on-demand or serverless options are the only correct choices.

614
MCQmedium

A DynamoDB-backed event processing system experiences throttling during a promotion. All events are written and read using the same partition key value (tenantId = "ACME"). The workload is time-ordered per tenant, and the application can tolerate slight reordering across partitions. Which design change will most directly increase throughput and reduce hot-partition throttling?

A.Increase the table's provisioned capacity (read/write units) to handle the promotion peak.
B.Change the partition key to include an additional sharding attribute derived from a hash of eventId.
C.Enable DAX caching for all reads but keep the same partition key and item layout.
D.Switch the table to eventually consistent reads for queries to lower read throttling.
AnswerB

When all traffic targets one partition key value, that partition becomes the bottleneck regardless of total table capacity. Adding a shard/salt attribute to the partition key (for example, tenantId + shardId where shardId = hash(eventId) mod N) spreads writes across multiple partition key values, increasing partition-level parallelism. Because the scenario allows slight reordering across partitions, losing strict single-partition time ordering is acceptable while improving throughput and reducing throttling.

Why this answer

Adding a sharding attribute derived from a hash of eventId allows writes and reads to be distributed across multiple partition keys, breaking the single hot partition caused by using tenantId='ACME' for all operations. DynamoDB's throughput is limited per partition, so distributing the load across many partitions directly reduces throttling without changing the application's tolerance for slight reordering.

Exam trap

The trap here is that candidates often assume increasing provisioned capacity (Option A) is the universal fix for throttling, but AWS specifically tests the understanding that DynamoDB's per-partition throughput limits require a sharding strategy to distribute load across partitions.

How to eliminate wrong answers

Option A is wrong because simply increasing provisioned capacity does not resolve the hot-partition issue; the single partition key (tenantId='ACME') still caps throughput at 3000 RCU/1000 WCU per partition, so throttling persists regardless of total table capacity. Option C is wrong because DAX caching only reduces read load on the table, but writes (which are the primary source of throttling during a promotion) still hit the same hot partition, and DAX does not help with write throttling. Option D is wrong because eventually consistent reads only reduce read costs and latency, but they do not address the root cause of throttling—the single partition bottleneck—and have no effect on write throttling.

615
Multi-Selectmedium

A company runs a web application on Amazon EC2 instances in multiple Availability Zones. The application uses an Application Load Balancer and an Auto Scaling group. The company wants to reduce cost while maintaining high availability. The workload is steady and predictable, and the instances run continuously. Which two actions will reduce cost? (Choose two.)

Select 2 answers
A.Move the application to a single Availability Zone to reduce data transfer charges.
B.Right-size the instances based on CloudWatch metrics to eliminate over-provisioned capacity.
C.Enable termination protection on all instances to prevent accidental deletion.
D.Replace the Application Load Balancer with a Network Load Balancer to reduce hourly charges.
E.Purchase Compute Savings Plans to cover the steady-state usage.
AnswersB, E

Right-sizing analyzes utilization metrics such as CPU, memory, and network to identify instances that are larger than needed. Moving to a smaller instance type reduces the hourly rate while still meeting performance requirements. For a steady workload, this is a reliable way to cut cost without sacrificing availability, provided the new size is validated under load.

Why this answer

For a steady, continuously running workload, the two most effective cost levers are purchasing commitments and right-sizing. Compute Savings Plans reduce the effective rate for the baseline usage, and right-sizing removes capacity that is not needed. High-availability features and load balancer choices should be preserved unless they are proven to be the primary cost driver.

Exam trap

The trap here is treating availability features or load balancer types as cost optimizations, when the real savings for steady workloads come from commitments and right-sizing.

616
MCQmedium

Account A hosts a role named AppReadRole. Account B needs to access it using STS AssumeRole. Account A’s role trust policy includes this condition: - StringEquals: { "sts:ExternalId": "b-7f9a" } When Account B runs: aws sts assume-role --role-arn arn:aws:iam::111111111111:role/AppReadRole --role-session-name test the call fails with: "AccessDenied: ExternalId mismatch". What should Account B change?

A.Provide the correct --external-id value (b-7f9a) in the AssumeRole call.
B.Add kms:Decrypt permissions to Account B’s IAM user because trust policy failures are KMS related.
C.Remove the ExternalId condition from the trust policy so any caller can assume the role.
D.Use AssumeRoleWithSAML instead of AssumeRole so ExternalId is not required.
AnswerA

The trust policy on the role includes a condition that requires sts:ExternalId to equal b-7f9a. When Account B's IAM user calls AssumeRole without an external ID or with an incorrect one, the condition fails and STS denies the request before any temporary credentials are issued. Supplying the matching --external-id b-7f9a in the CLI call satisfies the condition, allowing the AssumeRole to succeed and returning a temporary credential set.

Why this answer

The error 'AccessDenied: ExternalId mismatch' occurs because the trust policy on Account A's role requires an `sts:ExternalId` condition with the value `b-7f9a`, but Account B's `aws sts assume-role` command did not include the `--external-id` parameter. By providing the correct `--external-id b-7f9a` in the call, Account B satisfies the condition, allowing the role assumption to succeed. This is a standard security mechanism to prevent the confused deputy problem.

Exam trap

The trap here is that candidates may think the error is due to missing permissions (like KMS) or that changing the API method (SAML) avoids the condition, when in fact the fix is simply to include the required `--external-id` parameter in the AssumeRole call.

How to eliminate wrong answers

Option B is wrong because the error is explicitly about an ExternalId mismatch, not a KMS permissions issue; KMS is unrelated to STS AssumeRole trust policy conditions. Option C is wrong because while removing the condition would technically allow the call, it weakens security and is not the minimal change required to fix the mismatch error. Option D is wrong because AssumeRoleWithSAML does not bypass the ExternalId condition; the condition is evaluated regardless of the STS API used, and SAML-based calls have their own requirements.

617
MCQhard

Based on the exhibit, an EC2 application runs in private subnets with no NAT gateway and must retrieve a secret from AWS Secrets Manager. The secret uses a customer managed KMS key. Which change will allow the application to reach the service while keeping traffic off the internet?

A.Create an interface VPC endpoint for Secrets Manager and another interface VPC endpoint for KMS, and enable private DNS for both.
B.Create an S3 gateway endpoint for Secrets Manager and use the existing S3 gateway endpoint for both secret retrieval and KMS decryption.
C.Add a NAT gateway in a public subnet and route 0.0.0.0/0 from the private subnets to the NAT gateway.
D.Move the application into a public subnet so it can call the public Secrets Manager endpoint directly.
AnswerA

Secrets Manager is an interface endpoint service, and the customer managed KMS key means the application also needs private access to KMS for decrypt operations. Private DNS lets the SDK resolve standard service names to the VPC endpoints, keeping all traffic inside AWS private networking.

Why this answer

It creates interface VPC endpoints for both Secrets Manager and KMS, which allows the EC2 instance in the private subnet to securely access these services over the AWS network without traversing the internet. Enabling private DNS ensures that the standard service endpoints resolve to the private IP addresses of the VPC endpoints, eliminating the need for a NAT gateway or internet gateway.

Exam trap

The trap here is that candidates often assume a single endpoint type (like a gateway endpoint) can serve all AWS services, but Secrets Manager and KMS specifically require interface endpoints, and forgetting the KMS endpoint is a common oversight.

How to eliminate wrong answers

Option B is wrong because S3 gateway endpoints are designed for Amazon S3 only, not for Secrets Manager or KMS; Secrets Manager requires an interface endpoint (powered by AWS PrivateLink) and KMS also requires an interface endpoint or a separate connection. Option C is wrong because adding a NAT gateway would route traffic to the internet, which violates the requirement to keep traffic off the internet; the goal is to avoid internet-bound traffic entirely. Option D is wrong because moving the application to a public subnet would expose it to the internet, contradicting the requirement to keep traffic off the internet and potentially compromising security.

618
MCQmedium

A web application runs on an Amazon EC2 Auto Scaling group behind an Application Load Balancer (ALB). After each deployment, new instances take about 2 minutes to download artifacts and become ready to accept requests on the target port. In the last deployment, the ALB started marking targets unhealthy before the app was ready, and the Auto Scaling group then replaced those instances repeatedly, causing a prolonged outage. Which change best improves resilience during instance start-up without reducing actual availability once the application is healthy?

A.Increase the Auto Scaling group’s health check grace period so it exceeds the ~2-minute initialization time.
B.Add more subnets across additional Availability Zones to distribute the same instances more widely.
C.Switch the load balancer target type from instance targets to IP targets to avoid health check failures.
D.Reduce the ALB health check interval so unhealthy targets are removed faster.
AnswerA

A health check grace period prevents the Auto Scaling group from treating early health check failures as instance health problems. This avoids terminating instances before the application finishes initializing, which stops the restart/replace loop during deployments while still allowing normal health checks to apply once the app is ready.

Why this answer

The Auto Scaling group's health check grace period allows instances to initialize without being marked unhealthy by the ELB health checks. By setting this grace period to exceed the ~2-minute artifact download time, the ASG will not replace instances that are still starting up, preventing the cascade of terminations and redeployments that caused the outage. This directly addresses the root cause—premature health check failures—without changing the health check configuration or reducing availability once the app is ready.

Exam trap

The trap here is that candidates confuse the ALB health check interval or target type with the Auto Scaling group's lifecycle management, mistakenly thinking that changing how the ALB checks health (interval or target type) will fix the premature replacement, when the correct solution is to adjust the ASG's grace period to align with the application's startup time.

How to eliminate wrong answers

Option B is wrong because adding more subnets across additional Availability Zones distributes instances more widely for fault tolerance but does not prevent the ALB from marking starting instances as unhealthy, so it does not solve the premature replacement issue. Option C is wrong because switching from instance targets to IP targets changes how the ALB routes traffic but does not alter the health check logic or timing; the ALB will still mark the target as unhealthy if the health check fails during the initialization window. Option D is wrong because reducing the ALB health check interval causes unhealthy targets to be detected and removed faster, which would worsen the problem by accelerating the replacement cycle, not improving resilience during start-up.

619
MCQeasy

Account A hosts an IAM role that Account B developers must assume for a limited task. You want to require MFA for anyone assuming the role. Which trust policy condition most directly enforces that requirement for sts:AssumeRole?

A.Add a statement condition requiring "Bool": {"aws:MultiFactorAuthPresent": "true"} in the role trust policy.
B.Add a condition requiring "StringEquals": {"aws:PrincipalOrgID": "o-example"} without any MFA condition.
C.Add a statement that denies sts:AssumeRole when the requested role session name contains the text "dev".
D.Require HTTPS by setting a condition on "aws:SecureTransport": "true" in the trust policy.
AnswerA

The aws:MultiFactorAuthPresent context key is a boolean set by IAM to reflect whether the caller authenticated with an MFA device. Adding a condition requiring "Bool": {"aws:MultiFactorAuthPresent": "true"} in the role trust policy ensures that only principals who completed MFA during their own authentication can assume the role. If the caller did not use MFA, the key is either absent or false, and the trust policy evaluation fails, blocking sts:AssumeRole even if the identity is otherwise authorized.

Why this answer

The `aws:MultiFactorAuthPresent` condition key in the role trust policy directly checks whether the caller authenticated with a valid MFA device before calling `sts:AssumeRole`. When set to "true" with a Bool condition, it enforces that the session must have been established after MFA verification, which is the most direct and standard way to require MFA for role assumption.

Exam trap

The trap here is that candidates confuse transport-layer security (HTTPS) with authentication-layer MFA, thinking that requiring encrypted communication also enforces multi-factor authentication, but `aws:SecureTransport` only ensures the channel is encrypted, not that the caller proved possession of a second factor.

Why the other options are wrong

B

Option B does not enforce MFA; it restricts the role to principals from a specific AWS Organization, which is unrelated to MFA requirements.

C

This option denies AssumeRole based on the role session name containing 'dev', which does not enforce MFA. The question specifically asks for an MFA enforcement condition, not a naming restriction.

D

Requiring HTTPS (aws:SecureTransport) ensures encrypted transport but does not enforce MFA. The question specifically asks for MFA enforcement, so this condition does not address the requirement.

When would these options actually be correct?

B

This option would be correct if the question asked: 'Which condition restricts role assumption to principals belonging to a specific AWS Organization?'

C

In a scenario where an organization wants to restrict role assumption to only developers (e.g., by requiring the session name to include 'dev' as a naming convention), a condition like this could be used in the trust policy to allow or deny based on session name.

D

A question asks: 'You need to ensure that all API calls to assume a role are made over encrypted connections. Which condition should you add to the trust policy?' In that case, aws:SecureTransport would be the correct answer.

Why candidates pick the wrong answer

B

Candidates may confuse organizational controls with security requirements like MFA, or think that restricting by organization inherently includes MFA enforcement.

C

Candidates might think that restricting session names is a way to control access, but they overlook that the question explicitly requires MFA enforcement, not naming conventions.

D

Candidates may confuse security best practices (like using HTTPS) with specific MFA requirements, or think that transport encryption implies authentication strength.

620
MCQeasy

A customer-facing application has a relational data model and needs frequent complex queries (joins and aggregations), but it also experiences a significant read-heavy workload. Which design choice best improves read performance while keeping relational features?

A.Use DynamoDB with a single partition key and avoid indexes to keep writes simple.
B.Add read replicas to an RDS or Aurora cluster and keep the primary for writes.
C.Store the data in S3 and query it directly from the application without a database.
D.Switch the database to DynamoDB but keep using the same relational SQL queries and joins.
AnswerB

Read replicas offload read operations from the primary database instance, improving read throughput and reducing contention with writes. RDS/Aurora preserve relational capabilities like joins and SQL queries. This is a common and practical way to scale performance for read-heavy workloads without completely changing the data model.

Why this answer

Adding read replicas to an RDS or Aurora cluster offloads read traffic from the primary instance, improving read performance for complex queries (joins and aggregations) while preserving the relational data model. Aurora automatically scales read replicas and uses a shared storage volume, making this highly efficient for read-heavy workloads.

Exam trap

The trap here is that candidates often assume NoSQL databases like DynamoDB can handle relational queries if they simply 'switch' the database, ignoring that DynamoDB lacks native support for joins and complex aggregations, which are core to the relational data model described in the question.

How to eliminate wrong answers

Option A is wrong because DynamoDB with a single partition key and no indexes cannot efficiently support complex relational queries (joins and aggregations), and it sacrifices relational features. Option C is wrong because S3 is an object store, not a relational database; querying it directly for complex joins and aggregations is extremely slow and lacks transactional consistency. Option D is wrong because DynamoDB does not support SQL joins or complex relational queries; attempting to use the same relational SQL queries would fail or require significant application-level workarounds.

621
MCQmedium

A media company runs a two-tier web application in a VPC. The web tier is in public subnets behind an internet-facing Application Load Balancer, and the database tier is in private subnets running Amazon RDS. A security review found that the RDS security group allows traffic from 0.0.0.0/0 on port 3306. What is the MOST secure way to restrict database access to only the web tier?

A.Change the RDS security group inbound rule to allow port 3306 from the web tier's security group ID.
B.Create a network ACL on the database subnet that denies all traffic except port 3306 from the web subnet CIDR.
C.Change the RDS security group inbound rule to allow port 3306 from the VPC CIDR range.
D.Move the RDS instance into a public subnet and attach an Elastic IP, then restrict access by IP allowlist.
AnswerA

Referencing the web tier's security group as the source allows traffic only from instances that carry that group, regardless of their IP addresses. This is the recommended pattern because it survives scaling and IP changes and removes the wide-open CIDR. It directly closes the exposure while preserving connectivity between the tiers, making it the most secure and maintainable fix.

Why this answer

Security group referencing lets the database accept connections only from instances that belong to the web tier's security group, which is the tightest and most maintainable control. It removes the 0.0.0.0/0 exposure and automatically adapts as the web tier scales. VPC-wide CIDRs, public subnets, and network ACL CIDR rules are all broader or more fragile than a security group reference.

Exam trap

The trap here is choosing a CIDR-based rule for convenience, when security group referencing is both narrower and resilient to instance IP changes.

622
Multi-Selectmedium

An internal API is deployed in two AWS Regions behind separate Application Load Balancers. The company wants clients to use the primary Region when it is healthy and automatically switch to the secondary Region if the primary health check fails. Which two Route 53 record configurations are required? Select two.

Select 2 answers
A.Create a primary failover record that points to the primary ALB and associates a Route 53 health check.
B.Create a weighted record set that sends 50 percent of traffic to each Region.
C.Create a secondary failover record that points to the secondary ALB.
D.Create a latency-based record set so Route 53 always prefers the fastest Region.
E.Create a multivalue answer record to return both ALB addresses on each lookup.
AnswersA, C

A primary failover record designates the active endpoint in an active-passive configuration. Route 53 associates a health check with this record, and as long as the check returns healthy, every DNS response returns the primary ALB's IP address. When that health check fails, Route 53 stops serving the primary record and instead returns the secondary failover record, enabling automatic regional failover.

Why this answer

A primary failover record in Amazon Route 53 directs traffic to the primary ALB and is associated with a Route 53 health check. If the health check fails, Route 53 automatically fails over to the secondary failover record, ensuring high availability across Regions.

Exam trap

The trap here is that candidates often confuse failover routing with weighted or latency routing, assuming any health-aware routing provides automatic primary/secondary failover, but only failover records enforce a strict active-passive pattern.

623
MCQeasy

A company runs Amazon RDS for MySQL in a Multi-AZ configuration. If the primary database instance fails, what is the expected behavior?

A.The database remains unavailable until an administrator manually creates a new instance.
B.RDS automatically fails over to the standby instance in the same Region and keeps the same endpoint.
C.Traffic is routed to a read replica in another Region for immediate continuity.
D.The failed primary continues serving traffic while the standby synchronizes in the background.
AnswerB

Multi-AZ RDS is built for high availability. If the primary instance becomes unavailable, AWS automatically promotes the standby in the same Region and updates the DNS behind the database endpoint. Applications keep using the same connection string, so failover is largely transparent. This reduces downtime without requiring manual intervention or application changes.

Why this answer

Amazon RDS for MySQL in a Multi-AZ configuration automatically synchronously replicates data to a standby instance in a different Availability Zone. When the primary database instance fails, RDS automatically fails over to the standby instance, updating the DNS record for the same CNAME endpoint so that applications can resume operations without manual intervention. This ensures high availability with minimal downtime.

Exam trap

The trap here is that candidates often confuse Multi-AZ failover with read replicas, mistakenly thinking a read replica in another Region can serve as the failover target, whereas Multi-AZ uses a synchronous standby in the same Region.

Why the other options are wrong

A

Multi-AZ RDS automatically fails over to the standby in the same Region without manual intervention, so the database does not remain unavailable until an admin acts.

C

Read replicas in another Region are for read scaling and disaster recovery, not automatic failover. Multi-AZ failover uses a standby in the same Region, not a cross-Region read replica.

D

In a Multi-AZ RDS MySQL configuration, the standby instance is in a different Availability Zone, not in another Region, and the failed primary does not continue serving traffic; RDS automatically fails over to the standby.

When would these options actually be correct?

A

This would be correct if the question described a Single-AZ RDS deployment without Multi-AZ or automated failover, where an administrator must manually restore from a snapshot or create a new instance.

C

If the question were about disaster recovery using Amazon RDS cross-Region read replicas and asked how to promote a read replica to a primary for continuity after a regional failure, then traffic could be routed to the promoted read replica.

D

This option would be correct if the question described a scenario where RDS is configured with a Multi-Region read replica and the primary fails, but the application is designed to use the read replica for reads only, not for failover. However, for write continuity, a manual promotion would be needed.

Why candidates pick the wrong answer

A

Candidates may confuse Multi-AZ with Single-AZ deployments or assume that any database failure requires manual recovery, overlooking RDS's automated failover feature.

C

Candidates may confuse read replicas with Multi-AZ standby instances, assuming any replica can serve as a failover target, and may overestimate the automatic failover capabilities of cross-Region replicas.

D

Candidates may confuse Multi-AZ failover with read replica promotion, or incorrectly assume that the primary can still serve traffic during a failure if the standby is synchronizing.

624
MCQmedium

A static web application uses CloudFront with an S3 origin for assets (JavaScript, CSS, images). After deploying a new frontend build, the CloudFront cache hit ratio dropped significantly because the S3 origin receives many repeated requests for the same assets. The team notices that requests now include the Authorization header in asset requests. Which change is most likely to restore cache efficiency and reduce origin request costs?

A.Keep the Authorization header but increase the cache TTL to 1 year to reduce revalidation frequency.
B.Update the CloudFront cache policy so that Authorization is excluded from the cache key for static asset paths.
C.Remove CloudFront and serve assets directly from the S3 website endpoint to reduce CloudFront charges.
D.Switch the S3 origin from private access to public access so CloudFront can cache assets more effectively.
AnswerB

When Authorization is part of the cache key, each unique token can create separate cache entries, lowering the cache hit ratio and increasing origin requests. Excluding Authorization from the cache key (and typically from the origin request policy for static assets) allows caching to be based on the URL path/query string, improving hit ratio and reducing S3 origin load.

Why this answer

The drop in cache hit ratio is caused by the Authorization header being included in asset requests, which makes each request unique from CloudFront's perspective, preventing cache reuse. By updating the CloudFront cache policy to exclude the Authorization header from the cache key for static asset paths, CloudFront can treat identical asset requests as cache hits, restoring cache efficiency and reducing origin load.

Exam trap

The trap here is that candidates may assume increasing TTL or making the origin public solves caching issues, but the real problem is the cache key variation caused by the Authorization header, which must be explicitly excluded from the cache policy for static content.

How to eliminate wrong answers

Option A is wrong because increasing the TTL to 1 year does not address the root cause—the Authorization header still varies the cache key, so requests will continue to miss cache and revalidate unnecessarily. Option C is wrong because removing CloudFront and serving assets directly from the S3 website endpoint would eliminate caching entirely, increasing origin request costs and latency, not reducing them. Option D is wrong because switching the S3 origin from private to public access does not affect CloudFront's ability to cache; the cache key issue with the Authorization header remains, and public access introduces security risks without solving the problem.

625
MCQmedium

A dev sandbox has unpredictable DynamoDB traffic with long idle periods and occasional spikes. Which capacity mode should minimize operational overhead and avoid paying for idle provisioned capacity? The design must avoid adding custom operational scripts.

A.Reserved capacity for maximum daily traffic
B.Provisioned capacity set for peak traffic
C.DynamoDB on-demand capacity mode
D.Global tables in every Region
AnswerC

On-demand capacity mode bills per actual read and write request, using per-million request units, and automatically scales from zero to whatever the workload demands without capacity planning or throttling caused by forecasting errors. Idle sessions cost nothing beyond storage, and sudden traffic bursts are absorbed without needing to adjust provisioned units. For a sandbox with unpredictable traffic and long idle phases, on-demand avoids overpaying for unused capacity while providing exactly the elasticity this workload requires.

Why this answer

DynamoDB on-demand capacity mode is ideal for unpredictable workloads with long idle periods and occasional spikes because it automatically scales to handle traffic without requiring any capacity planning or provisioning. You pay only for the reads and writes you actually perform, eliminating the cost of idle provisioned capacity and the operational overhead of managing scaling scripts or alarms.

Exam trap

The trap here is that candidates may confuse 'reserved capacity' with DynamoDB's reserved capacity pricing model (which is actually a commitment discount for provisioned mode) or assume that provisioned capacity with auto-scaling is sufficient, but auto-scaling still requires setting minimum and maximum values and can incur costs for idle provisioned capacity during low-traffic periods.

How to eliminate wrong answers

Option A is wrong because Reserved capacity is not a DynamoDB pricing model; it applies to services like EC2 RIs or Aurora, and even if interpreted as provisioned capacity, it would require estimating peak traffic and paying for idle time. Option B is wrong because Provisioned capacity set for peak traffic would incur costs for unused capacity during idle periods and would require manual scaling or custom scripts to adjust capacity, violating the requirement to avoid custom operational scripts. Option D is wrong because Global tables replicate data across multiple Regions for disaster recovery or low-latency global access, which adds complexity and cost without addressing the core issue of unpredictable traffic and idle capacity waste.

626
MCQmedium

A SaaS platform plans to run in two AWS Regions for lower latency. The team wants to enable active-active writes (both regions accept updates) to avoid failover downtime. However, the business requires strong consistency for order status transitions (for example, only one transition from “Paid” to “Shipped” must be allowed). Which statement is the best architectural choice to meet the consistency requirement?

A.Use active-active writes only when the workload tolerates eventual consistency; for strongly consistent transitions, use a single-writer pattern with failover (active-passive/pilot light).
B.Active-active writes always provide strong consistency because AWS replicates data across Regions automatically and immediately.
C.Active-active writes can be used safely by simply enabling retries and expecting the application to resolve conflicts without coordination.
D.To ensure strong consistency, run both Regions with different IAM roles and block cross-Region writes at the API layer only.
AnswerA

Correct. Active-active multi-Region writes rely on asynchronous cross-Region replication (e.g., DynamoDB Global Tables or Aurora Global Database), so a write in one Region is not immediately visible in the other; concurrent updates to the same item can conflict and require last-write-wins or custom conflict resolution, which is inherently eventually consistent. Strongly consistent transitions, such as a transaction that must read its own write or maintain a linearizable order, cannot be guaranteed with multiple concurrent writers. A single-writer pattern (active-passive or pilot light) ensures only one Region accepts writes at a time; failover transfers write authority to the secondary Region, preserving strong consistency because there is never more than one authoritative writer.

Why this answer

Active-active writes across AWS Regions cannot guarantee strong consistency due to the inherent latency and lack of synchronous replication between Regions. For order status transitions that require exactly-once semantics (e.g., only one transition from 'Paid' to 'Shipped'), a single-writer pattern (active-passive or pilot light) ensures that only one Region accepts writes at a time, avoiding conflicts and maintaining a single source of truth. AWS services like DynamoDB global tables offer eventual consistency for multi-region writes, while Aurora Global Database provides read replicas with failover but not active-active writes for strong consistency.

Exam trap

The trap here is that candidates assume AWS's global services (like DynamoDB global tables or Aurora Global Database) inherently provide strong consistency for multi-region writes, when in fact they are designed for eventual consistency and require careful trade-offs for strict ordering requirements.

Why the other options are wrong

B

AWS does not replicate data across Regions automatically or immediately; cross-Region replication is asynchronous, so active-active writes cannot guarantee strong consistency.

C

Active-active writes without coordination cannot guarantee strong consistency for order status transitions; retries and application-level conflict resolution are insufficient to prevent two regions from simultaneously accepting conflicting transitions (e.g., 'Paid' to 'Shipped' in both regions).

D

Blocking cross-Region writes at the API layer with different IAM roles does not prevent concurrent writes from being accepted in both regions before the API layer can reject them, so it cannot guarantee strong consistency for order status transitions.

When would these options actually be correct?

B

A question stating that the application uses a single AWS Region with multi-AZ deployment and requires strong consistency for all reads and writes, where synchronous replication within a Region can provide immediate consistency.

C

In a scenario where the workload can tolerate eventual consistency and the application is designed to handle conflicts (e.g., using last-writer-wins or CRDTs), and the business requirement is high availability with low latency rather than strong consistency for specific transitions.

D

If the question required preventing cross-Region writes for security reasons (e.g., to enforce data sovereignty) and consistency was not a concern, then using different IAM roles to block writes at the API layer would be a valid approach.

Why candidates pick the wrong answer

B

Candidates may mistakenly believe that AWS handles cross-Region replication synchronously, similar to within-Region services like DynamoDB global tables or Aurora Global Database, which are actually asynchronous.

C

Candidates may assume that retries and application logic can resolve all conflicts, underestimating the difficulty of ensuring strong consistency across regions without a centralized coordinator or locking mechanism.

D

Candidates may think that IAM-based API restrictions can enforce write ordering across regions, but IAM cannot coordinate the timing of concurrent requests in a distributed system.

627
MCQeasy

A consumer application reads from an Amazon SQS queue. Some messages have an invalid format and always fail processing. They are retried repeatedly and consume consumer capacity. What is the best way to prevent these "poison pill" messages from blocking normal processing?

A.Enable long polling and increase the maximum message retention to 30 days.
B.Configure a dead-letter queue (DLQ) with a redrive policy and a maxReceiveCount.
C.Switch the queue to FIFO and disable retries in the consumer code.
D.Delete the main queue and recreate it after every failure.
AnswerB

A DLQ with a redrive policy isolates poison-pill messages. After a message fails processing and is received more than maxReceiveCount times, SQS stops returning it to the main queue and moves it to the DLQ. Normal messages continue to be processed without repeatedly consuming consumer capacity.

Why this answer

A dead-letter queue (DLQ) with a redrive policy and a maxReceiveCount allows messages that repeatedly fail processing to be moved to a separate queue after a specified number of receive attempts. This prevents poison pill messages from being retried indefinitely, freeing consumer capacity for valid messages. Amazon SQS automatically redirects messages to the DLQ once the maxReceiveCount threshold is exceeded, ensuring normal processing is not blocked.

Exam trap

The trap here is that candidates may think increasing retention or polling settings will solve the problem, but they fail to recognize that only a DLQ with a redrive policy isolates repeatedly failing messages from consuming consumer capacity.

How to eliminate wrong answers

Option A is wrong because enabling long polling and increasing maximum message retention does not address the root cause of invalid messages; it only reduces empty responses and keeps messages longer, but poison pills will still be retried. Option C is wrong because switching to a FIFO queue does not prevent poison pills; FIFO ensures exactly-once processing but still retries failed messages, and disabling retries in consumer code would cause message loss without moving them to a DLQ. Option D is wrong because deleting and recreating the main queue after every failure is disruptive, loses all messages, and does not provide a systematic way to isolate or inspect poison pills.

628
MCQmedium

A Lambda function in Account A must upload reports to an S3 bucket in Account B. Security does not want long-lived access keys anywhere, and the access should be easy to revoke from Account B. Which approach is best?

A.Create an IAM role in Account B that Account A can assume through STS, then grant the role S3 permissions.
B.Create an IAM user in Account B and store its access keys in Lambda environment variables.
C.Attach a security group to the Lambda function that allows outbound traffic to the bucket.
D.Use AWS Organizations SCPs to grant the Lambda function permission to write to the bucket.
AnswerA

Cross-account role assumption with AWS STS is the standard way to grant temporary access without sharing long-lived credentials. By placing the permissions on a role in Account B and controlling the trust policy there, the bucket-owning account keeps central control and can revoke access by changing the trust relationship or permissions. The Lambda execution role in Account A assumes the role when needed and receives short-lived credentials only.

Why this answer

It uses cross-account IAM roles with AWS Security Token Service (STS) to grant temporary credentials to the Lambda function. This avoids long-lived access keys, and the permissions can be revoked immediately by modifying or deleting the role in Account B, meeting the security requirements.

Exam trap

The trap here is that candidates may confuse security groups (network-layer controls) with IAM policies (identity-based access), or mistakenly think SCPs can grant cross-account permissions when they only act as guardrails.

Why the other options are wrong

B

Option B uses long-lived access keys stored in Lambda environment variables, violating the security requirement to avoid long-lived credentials. Additionally, revoking access requires deleting or rotating the keys in Account B, which is less straightforward than removing a role trust policy.

C

Security groups control network traffic at the instance level, not S3 bucket access. They cannot grant or deny API-level permissions to write objects; S3 uses IAM policies, bucket policies, or ACLs for authorization.

D

SCPs are used to centrally control permissions for all accounts in an AWS Organization, not to grant cross-account access to a specific Lambda function. They can only deny or allow permissions at the account level, not to individual resources like a Lambda function.

When would these options actually be correct?

B

This option would be correct if the question stated that the Lambda function must use long-lived credentials (e.g., for legacy system compatibility) and the security requirement to avoid them was absent. It might also be acceptable in a single-account scenario where IAM users are the only option.

C

A question where a Lambda function needs to access an S3 bucket in the same account, and the concern is network-level restriction (e.g., VPC endpoint). Then attaching a security group to the Lambda function (via VPC configuration) to allow outbound traffic to the S3 VPC endpoint would be correct.

D

An SCP would be correct if the question asked how to prevent all accounts in an organization from writing to a specific S3 bucket, or to enforce a policy that restricts S3 bucket access across multiple accounts for compliance reasons.

Why candidates pick the wrong answer

B

Candidates may think storing keys in environment variables is acceptable because it avoids hardcoding, and they may overlook the 'no long-lived access keys' constraint. The simplicity of creating an IAM user and directly using its keys can seem easier than setting up cross-account roles.

C

Candidates may confuse network access control (security groups) with identity-based access control (IAM), thinking that allowing outbound traffic to the bucket's IP range is sufficient to grant write permissions.

D

Candidates may confuse SCPs with IAM policies, thinking they can grant fine-grained permissions to individual resources, or they may overestimate the scope of SCPs as a tool for cross-account access.

629
Multi-Selectmedium

A healthcare company stores patient records in an Amazon DynamoDB table. The table must be recoverable to any point within the last 35 days, and the data must remain available if an entire AWS Region becomes unavailable. Which two actions should a solutions architect take to meet these requirements? (Choose two.)

Select 2 answers
A.Enable point-in-time recovery on the DynamoDB table.
B.Enable DynamoDB Streams and archive the stream to an Amazon S3 bucket.
C.Configure on-demand backup and restore with a daily backup schedule.
D.Enable DynamoDB Accelerator (DAX) for the table.
E.Create a DynamoDB global table with replica tables in additional AWS Regions.
AnswersA, E

Point-in-time recovery continuously backs up the DynamoDB table and allows restoration to any second within the last 35 days, up to the configured retention period. This directly satisfies the requirement to recover patient records to any point within 35 days. It protects against accidental writes or deletes but does not provide cross-Region availability, so it must be combined with another action.

Why this answer

Point-in-time recovery provides continuous backups that allow restoration to any second within the last 35 days, meeting the recovery requirement. DynamoDB global tables replicate the table across multiple AWS Regions and support automatic failover, meeting the Region-level availability requirement. Together, these two features address both dimensions of the scenario.

Exam trap

The trap here is treating DynamoDB Streams or on-demand backups as substitutes for point-in-time recovery, when only PITR provides second-level restore within a 35-day window.

630
Multi-Selecthard

A private application in two private subnets must download objects from S3 and read parameters from Systems Manager Parameter Store without routing traffic through the public internet. Which two components should the architect use?

Select 2 answers
A.Interface VPC endpoint for Systems Manager
B.Internet gateway attached to the VPC
C.NAT gateway in each Availability Zone
D.Gateway VPC endpoint for Amazon S3
AnswersA, D

The interface VPC endpoint for Systems Manager is correct because SSM is a services that requires an interface endpoint powered by AWS PrivateLink. This creates elastic network interfaces (ENIs) with private IPs from the subnet that allow the private application to reach SSM without traversing the public internet, a NAT gateway, or an internet gateway. Traffic stays entirely within the AWS network, satisfying the private-only requirement.

Why this answer

Interface VPC endpoints (AWS PrivateLink) enable private connectivity to Systems Manager Parameter Store by creating an elastic network interface in the subnet with a private IP, allowing the application to read parameters without traversing the internet. Gateway VPC endpoints for S3 provide private access to S3 objects via route table entries, using the S3 public IP space but staying within the AWS network, avoiding the need for an internet gateway or NAT gateway.

Exam trap

The trap here is that candidates often confuse gateway endpoints (used for S3 and DynamoDB) with interface endpoints (used for most other AWS services), and may incorrectly assume a NAT gateway or internet gateway is needed for private subnet outbound traffic, ignoring that gateway endpoints work via route tables without public IPs.

631
MCQmedium

A company serves versioned images from S3 through CloudFront. After a release, CloudFront origin fetches increased sharply and the monthly CloudFront bill went up. They reviewed CloudFront logs and found that many requests include a query string parameter `reqId` that is unique per request (for example, `...?v=2026-04-01&reqId=...`). The team currently forwards all query strings to the cache key. What change is most likely to reduce origin fetches and cost while keeping the versioned images correct?

A.Update the CloudFront cache policy to ignore `reqId` and include only the stable `v` query string parameter in the cache key.
B.Lower the CloudFront minimum TTL to 0 seconds so cached objects revalidate more often, reducing origin fetch volume.
C.Set the S3 bucket to use compression and enable S3 Transfer Acceleration to reduce origin fetch charges.
D.Disable forwarding of the query string to the origin, but keep using the full query string (including `reqId`) in the cache key.
AnswerA

Because `reqId` is unique per request, including it in the cache key prevents cache reuse (each request maps to a different cache entry), resulting in frequent origin fetches. Excluding `reqId` and keeping only `v` allows many requests for the same version to share cached objects, reducing origin traffic and cost while preserving correct version behavior.

Why this answer

The `reqId` query string parameter is unique per request, which forces CloudFront to treat each request as a distinct cache object when all query strings are forwarded to the cache key. By configuring the cache policy to include only the stable `v` parameter (the version identifier) and ignore `reqId`, CloudFront can serve cached responses for all requests with the same `v` value, drastically reducing origin fetches and lowering costs. This approach preserves correct versioned image delivery because the `v` parameter still differentiates between image versions.

Exam trap

The trap here is that candidates may think forwarding all query strings is harmless or that lowering TTL helps reduce origin fetches, but the real issue is cache key fragmentation caused by unique parameters like `reqId`.

How to eliminate wrong answers

Option B is wrong because lowering the minimum TTL to 0 seconds would cause CloudFront to revalidate cached objects more frequently, increasing origin fetches and costs, which is the opposite of the desired outcome. Option C is wrong because enabling S3 Transfer Acceleration and compression reduces data transfer latency and size but does not address the root cause of excessive origin fetches caused by unique query strings in the cache key. Option D is wrong because disabling forwarding of the query string to the origin while keeping the full query string (including `reqId`) in the cache key would still create unique cache objects for each `reqId`, failing to reduce origin fetches.

632
MCQmedium

A mobile game backend uses Amazon Aurora. The workload has many short-lived database connections from Lambda functions, causing connection storms. What should be added?

A.An internet gateway
B.S3 Select
C.RDS Proxy
D.A larger Route 53 hosted zone
AnswerC

RDS Proxy pools and reuses database connections, so Lambda invocations share established connections instead of opening a new one each time. This directly absorbs the short-lived connection storms described in the stem, preventing Aurora from exhausting its connection limit.

Why this answer

RDS Proxy is the correct solution because it sits between Lambda functions and the Aurora database, pooling and reusing database connections. This prevents connection storms by reducing the overhead of establishing new connections for each short-lived Lambda invocation, and it also helps manage IAM authentication for Lambda functions without storing database credentials.

Exam trap

The trap here is that candidates may think scaling the database (e.g., increasing instance size) is the answer, but the question specifically targets connection management, not compute or storage capacity, and RDS Proxy is the AWS-managed service designed exactly for this use case.

How to eliminate wrong answers

Option A is wrong because an internet gateway is used to enable VPC-to-internet communication, not to manage database connection pooling or reduce connection storms. Option B is wrong because S3 Select is a service for retrieving subsets of data from objects in S3 using SQL-like expressions, and it has no role in database connection management. Option D is wrong because a larger Route 53 hosted zone increases the number of DNS records you can host but does not affect database connection handling or reduce connection storms.

633
Multi-Selecthard

A private application in two private subnets must download objects from S3 and read parameters from Systems Manager Parameter Store without routing traffic through the public internet. Which two components should the architect use? The design must avoid adding custom operational scripts.

Select 2 answers
A.Interface VPC endpoint for Systems Manager
B.Internet gateway attached to the VPC
C.NAT gateway in each Availability Zone
D.Gateway VPC endpoint for Amazon S3
AnswersA, D

An interface VPC endpoint provisions an elastic network interface with a private IP in each subnet, carrying AWS PrivateLink traffic to Systems Manager Parameter Store. This satisfies the stem's requirement to read parameters without public internet routing and without custom operational scripts.

Why this answer

Option A is correct because an Interface VPC endpoint (powered by AWS PrivateLink) for Systems Manager creates elastic network interfaces in the private subnets with private IP addresses, allowing the instances to call the ssmmessages, ec2messages, and ssm endpoints privately without traversing the public internet. Option D is correct because a Gateway VPC endpoint for Amazon S3 adds a route-table target that lets traffic to S3 stay on the AWS private network, so the private application can download objects without an internet or NAT path. Together these two endpoints satisfy the requirement to reach both S3 and Parameter Store privately and require no custom operational scripts.

Option B is not appropriate because an Internet gateway provides public internet connectivity and would expose or require public routing, which the scenario explicitly forbids. Option C is not appropriate because a NAT gateway enables outbound internet access (and incurs cost per AZ) rather than keeping traffic private, so it does not meet the no-public-internet requirement.

Exam trap

The trap here is that candidates often confuse Gateway VPC endpoints (used for S3 and DynamoDB) with Interface VPC endpoints (used for most other AWS services like Systems Manager), and may incorrectly assume a NAT gateway or internet gateway is needed for private subnet access to AWS services.

634
MCQmedium

A media processing workflow uses CloudWatch Logs heavily. Retaining all debug logs forever is increasing costs. What should be configured? The design must avoid adding custom operational scripts.

A.Route 53 health checks
B.CloudWatch Logs retention policies per log group
C.CloudWatch detailed monitoring on all instances
D.AWS Config aggregation
AnswerB

CloudWatch Logs retention policies are configured per log group and automatically delete log events older than the specified interval (e.g., 30 days to 10 years). For a media processing workflow generating heavy log traffic, setting an appropriate retention period directly reduces storage cost and ensures compliance with data-retention requirements. Without a policy, logs default to "Never Expire" and accrue cost indefinitely.

Why this answer

CloudWatch Logs retention policies allow you to set a time-based expiration (e.g., 30 days) on log groups, automatically deleting old log events. This directly reduces storage costs without requiring custom scripts, as the retention policy is a native CloudWatch Logs feature configured per log group.

Exam trap

The trap here is that candidates may confuse CloudWatch Logs retention policies with CloudWatch metrics retention or detailed monitoring, thinking that reducing metric granularity will lower log storage costs, when in fact log retention is a separate, per-log-group setting.

How to eliminate wrong answers

Option A is wrong because Route 53 health checks monitor endpoint availability and DNS routing, not log retention or cost optimization. Option C is wrong because CloudWatch detailed monitoring increases metric frequency (1-minute intervals) and incurs additional costs, but does not manage log retention or deletion. Option D is wrong because AWS Config aggregation centralizes resource configuration snapshots and compliance rules, not log lifecycle management.

635
MCQmedium

A risk simulation workload uses CloudWatch Logs heavily. Retaining all debug logs forever is increasing costs. What should be configured? The design must avoid adding custom operational scripts.

A.CloudWatch Logs retention policies per log group
B.AWS Config aggregation
C.CloudWatch detailed monitoring on all instances
D.Route 53 health checks
AnswerA

By default, CloudWatch Logs retains log events indefinitely, so a risk simulation workload generating heavy log volume will accumulate storage costs without limit. Configuring a retention policy per log group—such as 30 or 90 days—automatically expires and deletes log events after the specified period, controlling costs and meeting data lifecycle requirements. This is the correct, direct way to manage log storage for CloudWatch Logs.

Why this answer

CloudWatch Logs retention policies allow you to set per-log-group expiration rules (e.g., 30 days, 90 days) to automatically delete old log events, directly reducing storage costs without custom scripts. Since the workload uses CloudWatch Logs heavily and retains debug logs forever, configuring a retention policy on each log group is the simplest, most cost-effective solution that requires no operational overhead.

Exam trap

The trap here is that candidates may confuse cost optimization features (like retention policies) with monitoring or compliance tools (like AWS Config or detailed monitoring), assuming that any AWS service that 'monitors' can also reduce log storage costs.

How to eliminate wrong answers

Option B is wrong because AWS Config aggregation is used to consolidate configuration and compliance data from multiple accounts/regions, not to manage log retention or cost. Option C is wrong because CloudWatch detailed monitoring on EC2 instances collects metrics at 1-minute intervals (vs. 5-minute basic), which increases costs and does not affect log retention or deletion. Option D is wrong because Route 53 health checks monitor endpoint availability and DNS routing, not log storage or lifecycle management.

636
MCQhard

A media company serves a global audience from an Amazon S3 bucket in the us-east-1 Region. Users in Asia and Europe report high latency when downloading large video files directly from the bucket. The company wants to reduce download latency for these users without changing the application's bucket names or rewriting the application to use a different endpoint. Which solution should a solutions architect recommend?

A.Create an Amazon CloudFront distribution with the S3 bucket as the origin, and have users access the distribution domain name.
B.Move the S3 bucket to a Region closer to the majority of users.
C.Enable S3 Transfer Acceleration on the bucket and have clients use the accelerated endpoint.
D.Enable S3 Cross-Region Replication to buckets in Asia and Europe, and update the application to select the nearest bucket.
AnswerA

CloudFront caches objects at edge locations worldwide, so users in Asia and Europe retrieve content from a nearby edge instead of the us-east-1 bucket. This dramatically reduces download latency for large files and repeated requests. It is the canonical AWS solution for global content delivery from S3, and the application can keep using the same bucket while clients simply use the CloudFront distribution domain, satisfying the no-rewrite constraint.

Why this answer

CloudFront caches S3 content at edge locations close to viewers, so global users download from a nearby edge rather than the us-east-1 origin. This lowers latency for large media files and requires no change to the bucket or application code beyond pointing clients at the distribution domain. Transfer Acceleration targets faster transfers but not edge caching, replication forces endpoint selection changes, and relocating the bucket favors only one region.

Exam trap

The trap here is confusing S3 Transfer Acceleration, which speeds transfers through edge routing but does not cache content, with CloudFront, which actually caches objects at the edge.

637
MCQhard

A company is building a serverless application that processes messages from an Amazon SQS queue using AWS Lambda. The application must not lose messages and must handle occasional downstream failures gracefully. The Lambda function sometimes fails due to a transient error in a downstream service. The company wants to ensure that failed messages are retried and eventually processed, but also wants to avoid infinite retries that could block the queue. What should the company do?

A.Increase the Lambda function's timeout and memory allocation to reduce the chance of transient errors.
B.Configure a dead-letter queue (DLQ) on the source SQS queue and set the maximumReceiveCount to a reasonable value.
C.Set the SQS queue's visibility timeout to a very high value so that messages are not retried until the downstream service recovers.
D.Configure the Lambda function to write failed messages to an Amazon S3 bucket and delete them from the queue.
AnswerB

Setting a DLQ on the source queue and a maximumReceiveCount allows messages that repeatedly fail to be moved to a separate queue after a set number of attempts. This prevents infinite retries and preserves failed messages for later analysis or reprocessing. Lambda automatically returns messages to the queue if the function fails, so the receive count increments with each attempt.

Why this answer

Using a dead-letter queue on the source SQS queue with a maximumReceiveCount ensures that messages are retried a limited number of times and then moved to a DLQ for separate handling. This prevents poison messages from blocking the queue and allows for later analysis or reprocessing, meeting the requirement to avoid infinite retries while not losing messages.

Exam trap

The trap here is assuming that increasing Lambda timeout or visibility timeout solves the problem of persistent downstream failures, when in fact a DLQ is needed to prevent infinite retries.

638
MCQmedium

A healthcare analytics team runs a containerized reporting service on Amazon ECS with the Fargate launch type in a single Availability Zone. The service must remain available if one Availability Zone fails, and it must scale automatically based on CPU utilization. The tasks are stateless and write output to Amazon S3. Which configuration should a solutions architect implement?

A.Create an ECS service with a task count of two, place the tasks in subnets in two Availability Zones, register them with a target group behind an Application Load Balancer, and attach a target-tracking scaling policy based on ECSServiceAverageCPUUtilization.
B.Configure the ECS service with two tasks in the same subnet of one Availability Zone and add a Network Load Balancer with cross-zone load balancing enabled.
C.Use the EC2 launch type with an Auto Scaling group spanning two Availability Zones and a scheduled scaling policy that adds instances at peak hours.
D.Deploy the ECS service with a task count of one and enable deployment circuit breaker with rollback so that a failed task is automatically replaced.
AnswerA

Running at least two tasks in subnets across two Availability Zones removes the single-zone dependency, and the Application Load Balancer health checks route around a task in a failed zone. Target-tracking on ECSServiceAverageCPUUtilization is the native Application Auto Scaling mechanism for ECS services, so the service adds and removes tasks automatically as CPU load changes.

Why this answer

Resilience against an Availability Zone failure requires task capacity in more than one zone, and an Application Load Balancer with health checks directs traffic only to healthy tasks. Because the tasks are stateless and store output in S3, no shared state must be replicated. Application Auto Scaling with a target-tracking policy on ECSServiceAverageCPUUtilization then adjusts task count to match demand without manual intervention.

Exam trap

The trap here is confusing the deployment circuit breaker, which only handles failed rollouts, with genuine multi-Availability-Zone redundancy and demand-based scaling.

639
MCQmedium

A company is deploying a high-performance computing (HPC) cluster with 16 EC2 instances. The workload requires the lowest possible network latency and highest throughput between all nodes for tightly coupled parallel MPI computations. Which EC2 placement group type should a solutions architect recommend?

A.Cluster placement group
B.Partition placement group
C.Spread placement group
D.No placement group — use Auto Scaling across multiple AZs
AnswerA

A Cluster placement group is the definitive choice for tightly coupled HPC workloads because it co-locates instances in a single Availability Zone with dedicated, high-bandwidth, low-latency connectivity. This hardware-level proximity minimizes network jitter and switch hops, enabling the sub-microsecond latency required for MPI-style distributed computing. Additionally, Cluster placement groups fully support Elastic Fabric Adapter (EFA), which bypasses the OS kernel to deliver near-bare-metal performance for tightly coupled parallel jobs, making it the standard placement strategy for HPC clusters.

Why this answer

Cluster placement groups pack instances physically close together within a single Availability Zone, providing the lowest possible network latency and highest network throughput between instances. They support enhanced networking (SR-IOV) and Elastic Fabric Adapter (EFA) for inter-node MPI communication.

Tightly coupled parallel HPC workloads require all nodes to communicate frequently with minimal latency. Cluster placement groups are specifically designed for this use case. The trade-off is all instances are in one AZ — if the AZ fails, the entire cluster is affected.

Exam trap

Spread and Partition placement groups improve availability by distributing instances across racks or partitions — they intentionally increase inter-node distance, which increases latency. For HPC requiring sub-microsecond inter-node communication, low latency trumps availability. Cluster PG = maximum performance in one AZ.

Spread PG = maximum isolation across racks.

Why the other options are wrong

B

Partition PGs distribute instances across separate hardware racks to reduce rack-failure impact. Instances in different partitions have higher inter-node latency. Designed for distributed databases (Hadoop, Cassandra, Kafka), not tightly coupled HPC.

C

Spread PGs place each instance on a distinct hardware rack for maximum isolation. Instances are intentionally spread further apart, increasing latency — the opposite of what HPC requires.

D

Multi-AZ Auto Scaling distributes instances across AZs for availability. Cross-AZ networking has higher latency than within-AZ. For HPC requiring sub-microsecond inter-node communication, all instances must be in the same AZ within a Cluster PG.

640
MCQeasy

Based on the exhibit, the web tier becomes unavailable if us-west-2a has an outage. What is the best change to improve resilience with the least redesign?

A.Increase the Auto Scaling group desired capacity from 2 to 3 in the same subnet.
B.Attach the Application Load Balancer and Auto Scaling group to subnets in a second Availability Zone.
C.Replace the Application Load Balancer with a Network Load Balancer.
D.Increase the health check grace period so instances stay registered longer.
AnswerB

Spanning the load balancer and Auto Scaling group across at least two Availability Zones removes the single-AZ dependency shown in the exhibit. If us-west-2a fails, the remaining AZ can continue serving traffic and Auto Scaling can replace unhealthy instances there. This is the smallest architectural change that directly improves availability.

Why this answer

The web tier is currently deployed in a single Availability Zone (us-west-2a), so an outage of that AZ makes the entire tier unavailable. By attaching the Application Load Balancer and Auto Scaling group to subnets in a second Availability Zone, the application can continue serving traffic from the healthy AZ, achieving high availability with minimal architectural changes. This is the standard AWS best practice for multi-AZ resilience.

Exam trap

The trap here is that candidates may think increasing instance count or changing load balancer type improves resilience, but the core issue is the single-AZ deployment, which only multi-AZ subnets can fix.

Why the other options are wrong

A

Increasing desired capacity in the same subnet does not add fault tolerance across Availability Zones; the web tier remains vulnerable to a single AZ outage.

C

Replacing the Application Load Balancer with a Network Load Balancer does not address the single-Availability Zone failure; the web tier would still be in us-west-2a only, so an outage of that AZ would still cause unavailability.

D

Increasing the health check grace period only delays instance deregistration during an outage, but does not address the root cause: the web tier is in a single Availability Zone. Instances in us-west-2a will still become unhealthy and eventually be terminated, causing unavailability.

When would these options actually be correct?

A

If the question asked how to handle increased traffic without changing AZ architecture, increasing desired capacity would be correct to distribute more instances within the existing subnet.

C

If the question required handling sudden traffic spikes with minimal latency and the application could tolerate connection draining, a Network Load Balancer might be chosen for its higher throughput and static IP support, but only if the architecture already spans multiple AZs.

D

A question where instances are being prematurely terminated due to transient health check failures (e.g., brief CPU spikes) and the goal is to avoid unnecessary instance replacement without changing architecture. The correct answer would be to increase the grace period to allow recovery.

Why candidates pick the wrong answer

A

Candidates may think more instances automatically improve resilience, overlooking that all instances are in the same AZ and thus share the same failure domain.

C

Candidates may think a Network Load Balancer is inherently more resilient, but resilience comes from multi-AZ deployment, not the load balancer type.

D

Candidates may think that giving instances more time to recover will prevent them from being marked unhealthy during a short outage, but they overlook that a full AZ outage is not recoverable by extending the grace period.

641
MCQmedium

You serve private reports stored in an S3 bucket through CloudFront. After a recent change, users report that they can access the S3 object URLs directly (bypassing CloudFront), which violates your design. You want to ensure S3 objects are readable only through CloudFront using Origin Access Control (OAC), even if someone guesses the S3 URL. Which update best enforces this at the S3 bucket level?

A.Add a bucket policy Allow for s3:GetObject only when the principal is cloudfront.amazonaws.com and aws:SourceArn matches your CloudFront distribution ARN, while blocking public access.
B.Enable an S3 bucket lifecycle policy to transition objects to Glacier, so public S3 URLs become inaccessible.
C.Rely only on CloudFront signed URLs validation; do not change the S3 bucket policy.
D.Add a WAF rule on CloudFront to block requests that contain "amazonaws.com" in the URL path.
AnswerA

CloudFront Origin Access Control (OAC) uses a service principal (cloudfront.amazonaws.com) to sign requests to S3. By adding a bucket policy that grants s3:GetObject only when the principal is cloudfront.amazonaws.com and aws:SourceArn matches your CloudFront distribution ARN, you ensure only that distribution can read objects. Blocking all public access with the S3 Block Public Access setting removes any other path, so a user who discovers the raw S3 URL cannot retrieve the object without going through CloudFront.

Why this answer

It uses an S3 bucket policy that grants s3:GetObject access only when the principal is cloudfront.amazonaws.com and the aws:SourceArn matches the CloudFront distribution ARN. This ensures that only CloudFront, using Origin Access Control (OAC), can retrieve objects, blocking direct S3 URL access even if the URL is guessed. Blocking public access at the bucket level further prevents any anonymous or public reads.

Exam trap

The trap here is that candidates may think CloudFront signed URLs alone are sufficient for security, but without a restrictive bucket policy, the S3 bucket remains publicly accessible, allowing direct URL access to bypass CloudFront.

How to eliminate wrong answers

Option B is wrong because a lifecycle policy to transition objects to Glacier does not prevent direct S3 URL access; it only changes storage class, and objects in Glacier are still accessible via S3 APIs if permissions allow. Option C is wrong because relying solely on CloudFront signed URLs without updating the S3 bucket policy leaves the bucket publicly accessible, allowing direct S3 URL access to bypass CloudFront entirely. Option D is wrong because a WAF rule on CloudFront that blocks requests containing 'amazonaws.com' in the URL path would block legitimate CloudFront requests to the S3 origin, and it does not prevent direct S3 URL access which bypasses CloudFront altogether.

642
Multi-Selecthard

A log archive has old unattached EBS volumes and many stale snapshots. Which two actions reduce storage cost without affecting running instances?

Select 2 answers
A.Stop all EC2 instances in the account
B.Disable CloudTrail logging
C.Delete unattached EBS volumes after verifying they are no longer needed
D.Apply snapshot lifecycle policies to expire obsolete snapshots
AnswersC, D

Unattached EBS volumes continue to be billed for their provisioned capacity and IOPS even when no instance is using them; deletion is the only way to eliminate those charges. Before deleting, you must verify that the volume is not needed for data recovery, future instance launches, or as a source for new snapshots. In a log archive scenario, these volumes are prime candidates for deletion because they represent pure waste, and this action directly addresses the cost of the unattached volumes listed in the scenario.

Why this answer

Unattached EBS volumes incur storage costs even when not in use, as EBS pricing is based on provisioned capacity per GB-month. Deleting them after verifying they are no longer needed eliminates this cost without affecting running instances, since attached volumes are untouched. This directly addresses the question's requirement to reduce storage costs without impacting running workloads.

Exam trap

AWS often tests the misconception that stopping instances or disabling services like CloudTrail reduces storage costs, but the trap here is that only direct actions on the storage resources themselves (deleting volumes and expiring snapshots) affect EBS and snapshot billing.

643
MCQmedium

A deployment engineer created an IAM role for an automation workflow (AppDeployRole). The role has an attached identity policy that allows iam:CreateRole for specific resource ARNs. However, the role is also created with a permission boundary named DeployBoundary. The DeployBoundary policy currently does not include the iam:CreateRole action. During execution, the automation fails with AccessDenied for iam:CreateRole, even though the attached identity policy allows it. What is the best fix?

A.Edit AppDeployRole’s attached identity policy to add iam:CreateRole again; permission boundaries only apply when permissions are missing.
B.Update DeployBoundary to allow iam:CreateRole for only the required resource ARNs, following least privilege.
C.Remove the permission boundary from the role because permission boundaries are not enforced at runtime.
D.Encrypt the deployment artifacts with KMS so IAM denies become KMS authorization failures.
AnswerB

IAM permission boundaries define the maximum set of permissions the role can use. To permit iam:CreateRole, the DeployBoundary must explicitly allow iam:CreateRole (and scope it to the required resources). The attached identity policy alone is not sufficient when the boundary is more restrictive.

Why this answer

B is correct because when an IAM role has a permission boundary, the boundary defines the maximum permissions the role can have. Even if the identity-based policy allows iam:CreateRole, the effective permissions are the intersection of the identity policy and the permission boundary. Since DeployBoundary does not include iam:CreateRole, the action is denied.

Updating the boundary to allow iam:CreateRole for the required resource ARNs, following least privilege, grants the necessary permission while still constraining the role.

Exam trap

The trap here is that candidates often think permission boundaries are optional or only restrict when the identity policy is too permissive, but in reality they are an absolute limit that always reduces effective permissions, so even if the identity policy allows an action, the boundary can deny it.

How to eliminate wrong answers

Option A is wrong because permission boundaries are not a fallback that only apply when permissions are missing; they are an upper limit that always applies, and adding the action again to the identity policy does not override the boundary. Option C is wrong because permission boundaries are enforced at runtime; removing the boundary would bypass security controls and is not a best practice. Option D is wrong because encrypting deployment artifacts with KMS does not affect IAM authorization for iam:CreateRole; KMS handles encryption/decryption, not IAM policy evaluation.

644
MCQmedium

A healthcare company stores patient records in an Amazon S3 bucket. Compliance requires that every object be encrypted at rest with a key that the company fully controls, including the ability to rotate and immediately revoke access. The security team also needs a record of every time the key is used to decrypt an object. Which encryption configuration should the company implement?

A.Use SSE-C where the customer provides the encryption key on each request.
B.Use SSE-KMS with a customer managed key in AWS KMS.
C.Use SSE-S3 with the default AWS managed key for the bucket.
D.Use client-side encryption with an AWS managed key before uploading objects.
AnswerB

A customer managed key in AWS KMS gives the company full control over rotation, key policy, and disabling the key to revoke access. Every Encrypt and Decrypt call is logged to AWS CloudTrail, satisfying the audit requirement. This matches all stated compliance needs for controls and visibility over the encryption key.

Why this answer

The requirement is customer-controlled encryption with the ability to rotate and revoke keys and to audit every decryption. A customer managed key in AWS KMS provides editable key policies, scheduled or on-demand rotation, the ability to disable the key to cut off access, and CloudTrail logging of cryptographic operations. Other S3 encryption options either lack customer key control or lack a per-use audit trail.

Exam trap

The trap here is assuming that any S3 encryption option provides an auditable key usage trail and customer-controlled revocation, when only customer managed KMS keys deliver both.

645
MCQmedium

A company hosts an application on EC2 instances in private subnets. The instances must (1) read objects from Amazon S3 and (2) retrieve secrets from AWS Secrets Manager. The team currently sends all outbound traffic through a NAT gateway to reach both services. They want to reduce monthly cost while keeping traffic private (no internet egress) and without changing application logic. Which change is the most cost-effective?

A.Create a Gateway VPC endpoint for S3 and an Interface VPC endpoint for Secrets Manager, and ensure the subnet route tables / endpoint routing directs those service calls to the endpoints instead of the NAT gateway.
B.Keep the NAT gateway, but add AWS WAF rules to block non-service outbound requests to reduce NAT usage.
C.Disable IPv4 on the VPC subnets and rely on IPv6-only egress to reduce NAT gateway costs.
D.Replace the NAT gateway with a VPC firewall appliance instance to proxy outbound calls and reduce NAT fees.
AnswerA

This is the most cost-effective change because it removes the need to traverse the NAT gateway for those AWS service calls. S3 uses a Gateway VPC endpoint (route-table-based) for traffic to the S3 prefix list, so requests to S3 stay on the AWS network. Secrets Manager uses an Interface VPC endpoint (ENIs with private DNS), so requests to Secrets Manager stay private within the VPC/VPC endpoint network path. Because the application still calls the same AWS APIs, there is no logic change, and NAT data-processing charges drop to near zero for S3/Secrets Manager traffic.

Why this answer

Gateway VPC Endpoints for S3 and Interface VPC Endpoints for Secrets Manager allow private connectivity to these AWS services without traversing the internet or a NAT gateway. This eliminates NAT gateway hourly charges and data processing fees, reducing costs while keeping traffic within the AWS network. The application logic remains unchanged as the endpoints are accessed via the same DNS names, with route tables directing traffic to the endpoints instead of the NAT gateway.

Exam trap

The trap here is that candidates may assume NAT gateways are the only way to provide private subnet internet access, overlooking that VPC endpoints can provide private, cost-effective connectivity to specific AWS services without internet egress.

How to eliminate wrong answers

Option B is wrong because AWS WAF is a web application firewall for HTTP/HTTPS traffic, not a mechanism to reduce NAT gateway costs; it does not eliminate the NAT gateway's hourly and per-GB data processing fees. Option C is wrong because disabling IPv4 and relying on IPv6-only egress would require the application to use IPv6 addresses, which changes the application logic and may not be supported by all services; additionally, NAT gateways are not used for IPv6 traffic (egress-only internet gateways are used), so this does not address the cost of the NAT gateway for IPv4 traffic. Option D is wrong because replacing the NAT gateway with a VPC firewall appliance instance still incurs instance costs and management overhead, and it does not eliminate the need for internet egress to reach S3 and Secrets Manager unless endpoints are used; it is not more cost-effective than using VPC endpoints.

646
MCQhard

A company uses AWS Organizations with multiple accounts. A security engineer must ensure that no IAM user in any member account can create an access key for the root user or perform any action as the root user, even if an administrator in that account tries to allow it. What should the security engineer do?

A.Enable AWS IAM Access Analyzer in each member account and review findings for root user access keys.
B.Enable multi-factor authentication for the root user in every member account and require it for all IAM users through a password policy.
C.Configure an AWS Config managed rule in each account that checks for root user access keys and sends an Amazon SNS notification to the security team.
D.Attach a service control policy to the organization root or OUs that denies all actions when the principal is the root user, and denies iam:CreateAccessKey for the root user.
AnswerD

Service control policies set the maximum permissions for member accounts and apply to all principals, including the root user of member accounts. A deny for root user actions and for creating root access keys enforces the restriction even if a local administrator tries to grant it. This is the preventive, organization-wide control.

Why this answer

A service control policy attached to the organization root or the relevant OUs is evaluated for all member account principals, including the root user. Denying root user actions and iam:CreateAccessKey for the root user enforces the restriction regardless of local administrator intent. Detective tools like IAM Access Analyzer and AWS Config, and authentication controls like MFA, do not block the action.

Exam trap

The trap here is assuming that a local administrator in a member account can override an organization policy, when service control policies cap the maximum permissions for that account.

647
MCQeasy

You must ensure that all requests to an S3 bucket use TLS (HTTPS). Which S3 bucket policy approach best enforces this requirement for S3 access?

A.Allow all principals to GetObject when aws:SecureTransport is true
B.Use a policy statement that explicitly Denies any action when aws:SecureTransport is false
C.Deny requests only when the bucket name is not matched exactly in the request
D.Require that the requester uses SSE-KMS and reject requests without SSE-KMS configuration
AnswerB

A bucket policy statement with Effect = Deny and a condition aws:SecureTransport = false blocks non-HTTPS requests. Because explicit Deny overrides Allow during policy evaluation, this prevents access for any request that does not use TLS, even if other statements grant permissions.

Why this answer

The `aws:SecureTransport` condition key evaluates whether the request was sent using TLS. By explicitly denying all S3 actions when `aws:SecureTransport` is false, any HTTP request is rejected, ensuring only HTTPS requests succeed. This approach uses an explicit deny, which overrides any allow, making it the most secure and reliable method to enforce TLS.

Exam trap

The trap here is that candidates confuse encryption in transit (TLS/HTTPS) with encryption at rest (SSE-KMS or SSE-S3), leading them to select Option D, which does not address the transport security requirement.

How to eliminate wrong answers

Option A is wrong because an allow statement with `aws:SecureTransport: true` does not block HTTP requests; it only permits HTTPS, but any other policy that allows access (e.g., a public bucket policy) could still allow HTTP requests. Option C is wrong because matching the bucket name has no relation to TLS enforcement; it addresses routing or bucket identification, not transport security. Option D is wrong because requiring SSE-KMS enforces encryption at rest, not encryption in transit (TLS); requests without SSE-KMS could still be sent over HTTP, violating the requirement.

648
MCQeasy

A company keeps daily database backups in an S3 bucket. They may restore from backups during the first 30 days if there is an issue. After 30 days, backups are rarely restored, but must be retained for 2 years. Which lifecycle strategy most cost-effectively meets these requirements?

A.Delete backups after 30 days to avoid storage costs, since restores are rare.
B.Keep all backups in S3 Standard for the entire 2-year retention period.
C.Use an S3 lifecycle policy to keep backups in S3 Standard for 30 days, then transition them to S3 Glacier Deep Archive for the remainder of the 2-year retention period.
D.Move backups to S3 Glacier Deep Archive immediately after creation, even for the first 30 days.
AnswerC

This is correct because an S3 Lifecycle policy can automate the transition of objects after a specified number of days. Storing backups in S3 Standard for the first 30 days ensures rapid restoration during the period when failures are most likely to be detected, then transitioning to S3 Glacier Deep Archive for the remaining ~23 months meets the 2-year retention requirement at drastically lower storage cost. Lifecycle transitions are metadata operations, so they incur no retrieval charge, making this the most cost-effective and compliant approach.

Why this answer

It balances cost and compliance: backups are kept in S3 Standard for the first 30 days when restores are frequent, ensuring low-latency access, then transitioned to S3 Glacier Deep Archive for the remaining retention period. S3 Glacier Deep Archive offers the lowest storage cost (approximately $0.00099/GB/month) for long-term retention, and the lifecycle policy automates the transition without manual intervention. This approach minimizes storage costs while meeting the 2-year retention requirement.

Exam trap

The trap here is that candidates may assume immediate deletion (Option A) or immediate archiving (Option D) are acceptable, failing to recognize the dual requirement of frequent access in the first 30 days and long-term retention at minimal cost, which the lifecycle policy elegantly addresses.

Why the other options are wrong

A

This option fails to meet the requirement that backups must be retained for 2 years, as it deletes them after only 30 days.

B

Keeping all backups in S3 Standard for 2 years incurs high storage costs for data that is rarely accessed after 30 days, making it not cost-effective.

D

Moving backups to Glacier Deep Archive immediately after creation would incur retrieval costs and delays for restores during the first 30 days, when restores are common, making it less cost-effective and operationally unsuitable.

When would these options actually be correct?

A

If the requirement were to only retain backups for 30 days with no longer-term retention needed, then deleting after 30 days would be the most cost-effective strategy.

B

If the requirement was to restore backups instantly at any time during the 2-year period with no retrieval delays, and cost was not a primary concern, then keeping all data in S3 Standard would be correct.

D

If the requirement stated that backups are never restored within the first 30 days and must be retained for 2 years with the lowest possible storage cost, then moving them immediately to S3 Glacier Deep Archive would be the most cost-effective strategy.

Why candidates pick the wrong answer

A

Candidates may focus solely on cost savings and overlook the explicit 2-year retention requirement, assuming that rare restores justify immediate deletion.

B

Candidates may think S3 Standard is the default safe choice and overlook the cost savings of transitioning to lower-cost storage classes for rarely accessed data.

D

Candidates may assume that the cheapest storage class is always best, overlooking the need for quick and frequent access during the initial 30-day period.

649
MCQhard

A media company uses an Amazon CloudFront distribution to serve content from a private S3 bucket. The security team wants to ensure that users cannot bypass CloudFront and access the S3 bucket directly, and that only the distribution can read objects. Which configuration should be implemented?

A.Create an origin access control (OAC) for the distribution, update the bucket policy to allow the OAC, and block public access on the bucket.
B.Configure the S3 bucket as a website endpoint and use a bucket policy that allows only the CloudFront distribution's IP addresses.
C.Set the S3 bucket ACL to private and enable CloudFront signed URLs for all objects.
D.Create an origin access identity (OAI) for the distribution, update the bucket policy to allow the OAI, and block public access on the bucket.
AnswerA

Origin access control (OAC) is the current recommended method to secure S3 origins for CloudFront. It supports SSE-KMS, dynamic requests, and signed requests. Updating the bucket policy to allow the OAC and enabling S3 Block Public Access ensures direct access is denied and only CloudFront can read objects.

Why this answer

Origin access control (OAC) is the modern, recommended way to secure S3 origins for CloudFront. It allows the distribution to sign requests to S3, and the bucket policy can be scoped to the OAC. Blocking public access ensures users cannot bypass CloudFront.

This satisfies both requirements.

Exam trap

The trap here is selecting origin access identity (OAI) out of familiarity, when OAC is the current recommended feature that supports additional capabilities like SSE-KMS.

650
MCQeasy

Based on the exhibit, what should the architect recommend to reduce inter-node latency for this workload?

A.Use a spread placement group so each instance is placed on separate hardware.
B.Launch the instances in a cluster placement group within the same Availability Zone.
C.Move the instances into different Availability Zones to improve fault tolerance.
D.Use a partition placement group to balance traffic across partitions.
AnswerB

A cluster placement group places instances close together in a single Availability Zone, which is the best choice for workloads that exchange many small messages and need very low network latency. This design is common for tightly coupled compute, analytics, and HPC-style applications. Because the workload is not bandwidth-saturated but latency-sensitive, proximity matters more than broader distribution.

Why this answer

A cluster placement group provides the lowest possible latency and highest packet-per-second performance by ensuring instances are placed in close physical proximity within a single Availability Zone. This is ideal for tightly coupled, high-performance computing workloads that require low inter-node latency.

Exam trap

The trap here is that candidates confuse placement group types, assuming 'spread' or 'partition' can also reduce latency, when only cluster placement groups are designed for low-latency, high-throughput networking within a single AZ.

Why the other options are wrong

A

A spread placement group places instances on separate hardware to reduce correlated failures, but it does not minimize inter-node latency; in fact, it may increase latency due to physical separation.

C

Moving instances into different Availability Zones increases inter-node latency due to physical separation, which contradicts the goal of reducing latency.

D

A partition placement group spreads instances across logical partitions, each on separate racks, to reduce correlated failures for large distributed workloads like HDFS or Cassandra. It does not minimize inter-node latency, which requires the low-latency, high-bandwidth network of a cluster placement group within a single AZ.

When would these options actually be correct?

A

When the question emphasizes high availability and fault tolerance, such as avoiding simultaneous failures from a single rack or hardware, and latency is not the primary concern.

C

When the requirement is to maximize fault tolerance and high availability, such as for a multi-AZ deployment of a critical application that must survive an AZ failure, even at the cost of higher latency.

D

For a large-scale distributed data processing workload (e.g., Hadoop, Kafka) that needs to tolerate rack-level failures while maintaining high throughput, a partition placement group would be correct. The question would emphasize fault isolation across racks rather than minimizing latency.

Why candidates pick the wrong answer

A

Candidates may confuse 'separate hardware' with 'closer proximity,' assuming that spreading instances reduces network hops, but spread groups actually increase physical distance between instances.

C

Candidates may confuse fault tolerance with performance optimization, or assume that distributing across AZs always improves reliability without considering latency trade-offs.

D

Candidates may confuse partition placement groups with cluster placement groups, thinking that 'partition' implies grouping instances together for performance, or they may overvalue fault tolerance features without recognizing the primary latency requirement.

651
MCQhard

A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The design must avoid adding custom operational scripts.

A.Amazon EFS with mount targets in multiple Availability Zones
B.S3 mounted as a POSIX file system without a file gateway
C.Instance store volumes
D.An EBS volume attached to all instances
AnswerA

EFS mount targets in multiple Availability Zones give each Linux instance a local endpoint, so file storage stays reachable if one AZ fails. EBS cannot attach across AZs, and the multi-AZ design needs no custom scripts.

Why this answer

Amazon EFS provides a fully managed, POSIX-compliant NFSv4.1 shared file system that can be mounted concurrently across multiple Linux EC2 instances. By deploying mount targets in multiple Availability Zones, the file system remains accessible even if one AZ fails, satisfying the high-availability requirement without any custom scripts.

Exam trap

The trap here is that candidates may confuse EBS Multi-Attach (which has strict limitations and requires cluster-aware file systems) with a true shared file system, or assume that S3 with a FUSE mount is a viable POSIX alternative without considering the operational overhead and lack of native consistency.

How to eliminate wrong answers

Option B is wrong because mounting S3 as a POSIX file system (e.g., via s3fs-fuse) requires custom operational scripts and does not provide native POSIX semantics or strong consistency, making it unsuitable for shared file storage across AZs. Option C is wrong because instance store volumes are ephemeral, tied to a single EC2 instance, and cannot be shared across instances or survive AZ failures. Option D is wrong because a single EBS volume cannot be attached to multiple EC2 instances; it can only be attached to one instance at a time, and while Multi-Attach EBS exists, it is limited to specific instance types and does not provide a shared file system without additional cluster-aware software.

652
MCQmedium

A research team runs a latency-sensitive distributed training job on Amazon EC2. They deploy 80 identical nodes that exchange small messages frequently and need low network jitter. The job must run entirely within one Availability Zone. Which placement group strategy should a solutions architect use to maximize intra-cluster network performance?

A.Use a cluster placement group to keep all instances in close proximity within the same Availability Zone.
B.Use a spread placement group to distribute instances across distinct hardware to reduce jitter.
C.Use a partition placement group and place each node into its own partition for uniform latency.
D.Do not use a placement group; rely on the default EC2 scheduling to balance latency and availability.
AnswerA

A cluster placement group is optimized to place instances close together (for example, within the same rack/cluster) to reduce latency and jitter for traffic between the instances. Because the workload runs in a single Availability Zone, the cluster placement group aligns with the requirement for strong locality and low-jitter communication.

Why this answer

A cluster placement group is the correct choice because it places all 80 EC2 instances in close physical proximity within a single Availability Zone, ensuring low-latency, high-bandwidth network connections with minimal jitter. This placement group type is specifically designed for tightly coupled, latency-sensitive workloads like distributed training that require frequent, small message exchanges, as it leverages non-blocking, high-throughput networking between instances.

Exam trap

The trap here is that candidates often confuse spread placement groups (which reduce jitter by isolating hardware failures) with cluster placement groups (which reduce jitter by minimizing physical distance), not realizing that jitter in this context is caused by network hops, not hardware faults.

Why the other options are wrong

B

A spread placement group distributes instances across distinct hardware to reduce correlated failures, but it does not provide the low latency and high bandwidth needed for tightly coupled, latency-sensitive distributed training. It can actually increase network jitter due to greater physical separation.

D

Default EC2 scheduling does not guarantee low latency or low jitter because instances may be placed on different physical hardware, increasing network hops and variability, which is unacceptable for a latency-sensitive distributed training job requiring consistent performance.

When would these options actually be correct?

B

When the requirement is to maximize availability and fault tolerance by ensuring instances are on separate hardware, such as for a critical application that must survive hardware failures, and low latency is not the primary concern.

D

A solutions architect is deploying a web application across multiple Availability Zones for high availability, and the application can tolerate moderate network latency. In this case, default scheduling is sufficient and avoids placement group constraints that could limit instance availability.

Why candidates pick the wrong answer

B

Candidates may mistakenly think that spreading instances reduces jitter by avoiding resource contention, but for latency-sensitive workloads requiring frequent small messages, proximity is more important than hardware diversity.

D

Candidates may assume that default scheduling provides adequate performance for most workloads, underestimating the strict low-latency and low-jitter requirements of tightly coupled HPC or distributed training jobs.

653
MCQmedium

A Lambda function for a mobile banking backend needs to read a database password. The password must rotate automatically every 30 days and should not be stored in environment variables. Which service should be used?

A.An encrypted object in Amazon S3
B.AWS Secrets Manager with rotation enabled
C.AWS Systems Manager Parameter Store SecureString without automation
D.A KMS-encrypted Lambda environment variable
AnswerB

AWS Secrets Manager is a purpose-built service for storing and managing secrets such as database credentials, and with rotation enabled it automatically updates the secret on a defined schedule using a Lambda-based rotation function. It encrypts secrets with KMS keys, integrates directly with services like RDS and Lambda, and provides a caching layer to reduce latency and cost for high-frequency reads. This makes it the correct choice for a mobile banking backend that must rotate credentials regularly without custom orchestration.

Why this answer

AWS Secrets Manager is the correct choice because it provides built-in automatic rotation of secrets (e.g., database passwords) with a configurable rotation interval (e.g., 30 days). It integrates natively with AWS Lambda via the AWS SDK, allowing the function to retrieve the password at runtime without storing it in environment variables. Secrets Manager also encrypts secrets at rest using KMS and supports automatic rotation via a Lambda rotation function, meeting both the security and rotation requirements.

Exam trap

The trap here is that candidates often confuse AWS Systems Manager Parameter Store SecureString with Secrets Manager, but Parameter Store does not support automatic rotation without custom automation, making it unsuitable for a 30-day rotation requirement.

How to eliminate wrong answers

Option A is wrong because storing an encrypted object in Amazon S3 does not provide automatic rotation of the password; you would need to manually manage rotation and versioning, and the Lambda function would require additional logic to decrypt and rotate the secret. Option C is wrong because AWS Systems Manager Parameter Store SecureString without automation lacks built-in automatic rotation; you would need to implement custom rotation logic, and Parameter Store does not natively support scheduled rotation like Secrets Manager. Option D is wrong because a KMS-encrypted Lambda environment variable is static and cannot be rotated automatically; the password would remain the same until the Lambda function is redeployed, and environment variables are visible in the function configuration, violating the requirement to not store the password in environment variables.

654
MCQmedium

A game streaming service must use UDP for real-time gameplay traffic. For external firewall allowlisting, the service requires stable, static IP addresses. The TLS handshake must be handled end-to-end by the application servers (the load balancer must not terminate TLS). Which AWS load balancing option best fits these requirements?

A.Use a Network Load Balancer (NLB) with a UDP listener, configure the NLB to use Elastic IP addresses for static IPs, and use TCP listeners for TLS passthrough to the application servers.
B.Use an Application Load Balancer (ALB) with UDP listeners and configure TLS passthrough.
C.Use Amazon API Gateway with a WebSocket API and keepalive pings to provide UDP-like low-latency delivery.
D.Use a Classic Load Balancer and multiplex UDP over TCP to meet the UDP and low-latency requirements.
AnswerA

NLB supports UDP listeners and is designed for low-latency, high-performance networking. Associating Elastic IP addresses with the NLB provides stable public IP addresses for firewall allowlisting. For TLS passthrough, using a TCP listener keeps the TLS handshake and encryption between the client and the targets (no load balancer TLS termination).

Why this answer

A Network Load Balancer (NLB) supports UDP listeners, which are required for real-time gameplay traffic, and can be assigned Elastic IP addresses to provide stable, static IPs for firewall allowlisting. Additionally, NLB supports TCP listeners with TLS passthrough, meaning it forwards the encrypted traffic without terminating the TLS handshake, allowing the application servers to handle end-to-end encryption as required.

Exam trap

The trap here is that candidates may assume an ALB can handle UDP traffic because it supports WebSocket or HTTP/2, but ALB is strictly Layer 7 and only supports TCP-based protocols, while NLB is the correct choice for UDP and TLS passthrough with static IPs.

Why the other options are wrong

B

ALB does not support UDP listeners; it only supports HTTP, HTTPS, and WebSocket protocols. Additionally, ALB cannot provide static IP addresses or perform TLS passthrough without terminating TLS.

C

Amazon API Gateway WebSocket API does not support UDP traffic; it uses WebSocket protocol over TCP, and it cannot provide static IP addresses for firewall allowlisting. The requirement for UDP and static IPs makes this option invalid.

D

Classic Load Balancers do not support UDP listeners; they only handle TCP/SSL traffic. Multiplexing UDP over TCP would break real-time gameplay requirements by introducing TCP overhead and latency.

When would these options actually be correct?

B

When the requirement is for HTTP/HTTPS traffic with path-based routing, host-based routing, or WebSocket support, and TLS termination at the load balancer is acceptable. For example, a web application needing advanced routing and SSL offloading.

C

This option would be correct for a real-time chat or notification service that requires persistent bidirectional communication with low latency, where clients connect via WebSocket and the service does not need static IPs or UDP transport.

D

If the question required load balancing legacy HTTP/HTTPS traffic with basic round-robin routing and no need for UDP or static IPs, a Classic Load Balancer would be a valid, cost-effective option.

Why candidates pick the wrong answer

B

Candidates may mistakenly think ALB supports UDP because it handles WebSocket traffic, or they confuse ALB's TLS termination with the ability to pass through TLS without decryption.

C

Candidates may confuse WebSocket's low-latency, full-duplex communication with UDP's real-time capabilities, and overlook that WebSocket runs over TCP and lacks static IP support.

D

Candidates may think Classic Load Balancers can be adapted to support UDP via workarounds, or they confuse Classic Load Balancers with Network Load Balancers, assuming both support UDP.

655
MCQmedium

A company uses IAM permission boundaries to prevent developers from escalating privileges. The security team created a permission boundary that allows only read-only actions on most AWS services, but teams can still manage their own resources. A developer can create an IAM role with broad permissions, and the boundary does not appear to be restricting it. Which corrective action best aligns with how permission boundaries work?

A.Rely on an AWS-managed policy attached to the developer’s IAM user; permission boundaries only apply to users.
B.Ensure the role creation process sets the permission boundary on the new role, using the boundary’s ARN in the CreateRole call or role template.
C.Attach the permission boundary policy as an SCP in AWS Organizations so it automatically applies to all roles.
D.Grant the developer IAM permissions to add a “deny” statement to the boundary policy so the boundary blocks escalation.
AnswerB

Permission boundaries are evaluated based on the boundary attached to the principal/role being created or used. If a developer creates roles without specifying the boundary, the boundary won’t restrict the resulting permissions. Enforcing boundary attachment via role templates or required parameters ensures every created role is constrained.

Why this answer

Permission boundaries must be explicitly applied to a role during its creation (via the `CreateRole` API call or an infrastructure-as-code template). Without setting the boundary ARN, the role inherits no restriction, allowing the developer to create a role with broad permissions that bypasses the intended boundary. Option B correctly identifies that the role creation process must include the boundary ARN to enforce the limitation.

Exam trap

The trap here is that candidates assume permission boundaries are automatically inherited or enforced by default, when in fact they must be explicitly applied to each role during creation, and SCPs are often confused as a substitute for permission boundaries.

Why the other options are wrong

A

Permission boundaries apply to IAM roles and users, not just users. The developer can create a role without a boundary, so attaching a policy to the user does not restrict the role's permissions.

C

SCPs apply to all accounts in an AWS Organization but do not replace or enforce IAM permission boundaries on individual roles; permission boundaries must be explicitly set on each role during creation.

D

Permission boundaries cannot be modified by the user they restrict; only the boundary's creator (e.g., security team) can update it. Granting the developer permission to add a deny statement would violate the boundary's purpose and is not a valid corrective action.

When would these options actually be correct?

A

If the question asked how to restrict a developer's own permissions (not roles they create), attaching an AWS-managed policy to the user would be correct.

C

In a question where the goal is to enforce a maximum permission baseline across all accounts in an AWS Organization, and the requirement is to prevent any IAM entity from exceeding a defined set of actions, an SCP would be the correct answer.

D

In a scenario where a developer needs to implement additional restrictions on a role they manage, and the security team has delegated authority to modify a custom boundary policy for specific use cases, allowing the developer to add deny statements could be correct if explicitly authorized.

Why candidates pick the wrong answer

A

Candidates may think permission boundaries are only for users, or that a user's attached policy limits all actions they perform, including role creation.

C

Candidates may confuse SCPs with permission boundaries because both can restrict permissions, but they operate at different levels (account vs. entity) and have different enforcement mechanisms.

D

Candidates may think that adding a deny statement to the boundary policy would block privilege escalation, misunderstanding that permission boundaries are set by an admin and cannot be altered by the user they constrain.

656
MCQhard

Based on the exhibit, the security team needs to detect and alert on both successful and failed attempts to change S3 bucket policies and KMS key policies across the organization. Which solution best meets that requirement?

A.Enable an organization trail for management events in all regions and create an EventBridge rule that matches PutBucketPolicy and PutKeyPolicy, then send alerts to SNS.
B.Enable AWS Config in all accounts and use only a periodic compliance evaluation to alert when bucket or key policies drift.
C.Use IAM Access Analyzer because it continuously blocks policy changes that would expose the resources publicly.
D.Turn on S3 server access logging and KMS key rotation, because both services will capture policy modifications automatically.
AnswerA

CloudTrail management events record API activity, including failed attempts, and an organization trail provides coverage across accounts and Regions. EventBridge can react to those API calls in near real time and route notifications to SNS. This is the clean detective-control pattern for policy-change auditing.

Why this answer

AWS CloudTrail management events capture all API calls that modify S3 bucket policies (PutBucketPolicy) and KMS key policies (PutKeyPolicy). By enabling an organization trail for all regions, you centralize these events across the entire AWS Organization. An Amazon EventBridge rule can then filter for these specific API calls and send alerts via Amazon SNS, meeting the requirement to detect both successful and failed attempts.

Exam trap

The trap here is that candidates often confuse AWS Config's compliance checks or IAM Access Analyzer's policy analysis with real-time API call monitoring, failing to realize that only CloudTrail management events capture every attempt (including failures) to change policies.

How to eliminate wrong answers

Option B is wrong because AWS Config periodic compliance evaluations only check resource compliance at scheduled intervals, not in real-time, and they do not directly capture or alert on every API call attempt (including failed ones) to change policies. Option C is wrong because IAM Access Analyzer is designed to analyze resource-based policies for unintended public or cross-account access, not to block or alert on all policy change attempts; it does not continuously block changes or capture failed attempts. Option D is wrong because S3 server access logging logs object-level access requests, not management API calls like PutBucketPolicy, and KMS key rotation does not capture policy modifications; neither service logs policy change attempts.

657
MCQmedium

You have an S3 bucket that stores customer-specific private files. You want to serve these files through CloudFront, where clients must use signed cookies (or signed URLs) to access the content. In addition, you need to block common web exploits and rate-limit suspicious traffic at the edge. Which design best meets these requirements?

A.Keep the S3 bucket private, configure CloudFront with Origin Access Control so only CloudFront can access the origin, require signed cookies/URLs for viewers, and associate an AWS WAF web ACL with CloudFront for blocking and rate limiting.
B.Enable public read access on the S3 bucket and rely on WAF alone for authorization because WAF can validate signatures.
C.Configure CloudFront with signed URLs but do not change the S3 bucket access settings; leaving public access enabled is acceptable since CloudFront can filter traffic.
D.Use WAF at CloudFront but omit signed cookies/URLs because rate limiting and exploit blocking already provide access control for private files.
AnswerA

This ensures S3 remains non-public while CloudFront becomes the only origin access path using Origin Access Control. Signed cookies/URLs enforce authenticated authorization at the edge for each request. Attaching AWS WAF adds request inspection and protections like rate limiting and exploit blocking.

Why this answer

It combines a private S3 bucket with Origin Access Control (OAC) to ensure only CloudFront can access the origin, enforces signed cookies/URLs for viewer authentication, and uses AWS WAF at the edge to block common web exploits and rate-limit suspicious traffic. This layered approach provides both authorization (via signed requests) and security filtering (via WAF) at the CloudFront edge, meeting all requirements.

Exam trap

The trap here is that candidates often think WAF can handle authorization (like validating signed URLs) or that leaving the S3 bucket public is acceptable if CloudFront is used, but WAF cannot verify cryptographic signatures and a public bucket allows direct access bypassing CloudFront's authentication.

How to eliminate wrong answers

Option B is wrong because enabling public read access on the S3 bucket bypasses the need for signed cookies/URLs, and WAF cannot validate signatures—WAF inspects HTTP headers, URI paths, and IP addresses, but does not have the capability to verify CloudFront signed URL or signed cookie cryptographic signatures. Option C is wrong because leaving the S3 bucket publicly accessible defeats the purpose of using signed URLs; CloudFront does not filter traffic based on signed URLs at the origin level, so a public bucket would allow direct access to objects without authentication. Option D is wrong because omitting signed cookies/URLs means there is no mechanism to restrict access to authorized viewers only; WAF rate limiting and exploit blocking do not provide authentication or authorization for private content.

658
MCQeasy

A media company stores finalized video masters in an Amazon S3 bucket in the us-east-1 Region. Compliance requires that the objects be recoverable if they are accidentally deleted or overwritten for at least 90 days, and that no user, including administrators, be able to permanently erase them during that period. Which S3 feature should the solutions architect enable?

A.S3 Versioning with a lifecycle rule that transitions noncurrent versions to S3 Glacier Deep Archive after 30 days.
B.S3 Object Lock in governance mode with a retention period of 90 days on the bucket.
C.S3 Object Lock in compliance mode with a retention period of 90 days on the bucket.
D.S3 Cross-Region Replication to a bucket in another Region with versioning enabled on both buckets.
AnswerC

S3 Object Lock in compliance mode prevents any user, including the root account, from deleting or overwriting a protected object version until the retention period expires. A 90-day retention period matches the compliance window exactly, and the protection cannot be shortened or bypassed, which is what the requirement demands.

Why this answer

The requirement is legal-hold-style immutability that even administrators cannot override. S3 Object Lock in compliance mode enforces exactly that: protected object versions cannot be deleted or overwritten by any identity until the retention period lapses, and the retention cannot be shortened. Governance mode and replication both leave a path for privileged users to destroy the data.

Exam trap

The trap here is treating governance mode and compliance mode as interchangeable, when only compliance mode removes the ability of privileged users to bypass retention.

659
MCQhard

Based on the exhibit, the database is manually promoted during an Availability Zone failure and the application outage lasts longer than the target. What change best improves resilience with the least operational intervention?

A.Keep the read replica and automate promotion with a runbook after CloudWatch alarms fire.
B.Convert the database to an RDS Multi-AZ deployment so a synchronous standby can fail over automatically.
C.Use a cross-Region read replica so promotion happens faster during an AZ failure.
D.Increase the application retry count and keep the current database design.
AnswerB

Multi-AZ is designed for automatic failover within the same Region and maintains a synchronous standby for high availability. The exhibit shows that the current read replica requires manual promotion and produces an outage longer than the target. Switching to Multi-AZ removes the manual step and aligns the database layer with the desired recovery time.

Why this answer

B is correct because RDS Multi-AZ automatically synchronously replicates data to a standby in a different Availability Zone and triggers an automatic failover with zero manual intervention when an AZ failure occurs. This directly addresses the requirement to improve resilience while minimizing operational effort, as the failover is handled by AWS without any runbook execution or manual promotion.

Exam trap

The trap here is that candidates often confuse read replicas (designed for read scaling and manual promotion) with Multi-AZ deployments (designed for automatic failover), and incorrectly assume that automating a runbook for read replica promotion is equivalent to the native automatic failover of Multi-AZ.

Why the other options are wrong

A

Manual promotion via runbook after CloudWatch alarms still requires human intervention, which does not meet the goal of 'least operational intervention' and will not reduce outage duration as effectively as automatic failover.

C

Cross-Region read replicas are asynchronous and do not support automatic failover; promoting them requires manual intervention and takes longer than Multi-AZ failover, so they do not improve resilience with the least operational intervention during an AZ failure.

D

Increasing the application retry count does not address the root cause of the outage (AZ failure) and does not improve resilience; it only masks the symptom and may lead to degraded user experience or timeouts.

When would these options actually be correct?

A

This option would be correct in a scenario where the database must remain in a single-AZ configuration due to cost constraints or compliance, but the organization can tolerate a slightly longer outage and wants to minimize manual steps by automating the promotion process with a runbook triggered by CloudWatch alarms.

C

This option would be correct if the question required resilience against a Region-wide outage (not just an AZ failure) and the application could tolerate eventual consistency, as cross-Region replicas provide disaster recovery across geographic regions.

D

In a scenario where the database is already resilient (e.g., Multi-AZ) but transient network blips cause brief connection drops, increasing the retry count can help the application recover without manual intervention.

Why candidates pick the wrong answer

A

Candidates may think automating the runbook reduces operational effort, but they overlook that manual promotion still requires human action, which is slower and more error-prone than automatic failover.

C

Candidates may think cross-Region replicas offer faster promotion because they are in a different location, but they overlook the asynchronous replication lag and lack of automatic failover, which actually increases outage duration.

D

Candidates may think that retries can compensate for any failure, underestimating the need for infrastructure-level high availability and overestimating application-level fault tolerance.

660
MCQhard

Based on the exhibit, which design change is the best way to reduce the observed read latency for this DynamoDB-backed service?

A.Add a DynamoDB Accelerator (DAX) cluster in front of the table and send repeated read traffic through it.
B.Increase the on-demand table limits so DynamoDB can automatically absorb more traffic.
C.Create a global secondary index on tenantId to distribute the load across more partitions.
D.Move the dashboard data into S3 and use Lambda functions to read it on demand.
AnswerA

DAX is designed to accelerate repeated eventually consistent reads from DynamoDB by caching hot items in memory. The exhibit shows one tenant driving most of the reads and the same dashboard items being requested repeatedly within a short window, which is an excellent fit for DAX. It reduces latency and offloads the hot key without requiring a schema redesign.

Why this answer

Adding a DynamoDB Accelerator (DAX) cluster in front of the table reduces read latency by providing an in-memory cache for repeated read traffic. DAX delivers microsecond response times for eventually consistent reads, which directly addresses the observed latency issue without requiring application-level caching or table redesign.

Exam trap

The trap here is that candidates often assume increasing capacity limits (Option B) or adding indexes (Option C) will solve latency issues, but they fail to recognize that latency is a caching problem, not a throughput or partitioning problem, and that DAX is the AWS-native solution for DynamoDB read-heavy workloads with repeated access patterns.

Why the other options are wrong

B

Increasing on-demand table limits does not reduce read latency; it only prevents throttling. The observed latency is likely due to repeated reads of the same hot data, which DAX caching addresses directly.

C

A GSI on tenantId does not reduce read latency for repeated reads of the same data; it only helps with query flexibility. The observed latency is likely due to hot partitions or throttling, which DAX's caching directly addresses.

D

Moving dashboard data to S3 and using Lambda to read it on demand would likely increase read latency due to cold starts and S3's higher latency for small, frequent reads compared to DAX. It also adds complexity and cost without addressing the root cause of high read latency on DynamoDB.

When would these options actually be correct?

B

This option would be correct if the question described a scenario where the application is experiencing ProvisionedThroughputExceededException errors due to insufficient read capacity, and the workload is unpredictable, making on-demand capacity the appropriate solution.

C

If the question described a scenario where read traffic is evenly distributed across many distinct tenantId values and the goal is to improve query performance by avoiding full table scans, then adding a GSI on tenantId would be correct.

D

This option would be correct if the question involved large, infrequently accessed dashboard data (e.g., historical reports) where cost savings from S3's lower storage cost outweigh latency concerns, and the Lambda function can be optimized with provisioned concurrency to minimize cold starts.

Why candidates pick the wrong answer

B

Candidates may assume that increasing capacity limits will improve performance, confusing throughput with latency, or they may not understand that DynamoDB's on-demand scaling handles capacity but not caching or hot-key issues.

C

Candidates may think that distributing load across partitions via a GSI will reduce latency, but they overlook that GSIs don't cache data and that the bottleneck is likely from repeated reads of the same items, not partition distribution.

D

Candidates may think S3 is always faster or cheaper for any data, or they might overestimate Lambda's ability to reduce latency without considering cold starts and network overhead.

661
MCQeasy

An application runs on an EC2 Auto Scaling group. Over the last month, CPU utilization averaged 8% with no sustained memory pressure, and response times are stable. The team wants to lower monthly cost without changing the application. What is the most appropriate next step for cost optimization?

A.Evaluate a smaller EC2 instance type (via the Auto Scaling launch template/configuration) for the group and validate performance metrics after the change.
B.Increase desired capacity to 2x so utilization increases and instances become “more efficient.”
C.Disable Auto Scaling so the group never scales down to preserve baseline performance.
D.Switch the workload to Spot instances immediately to avoid On-Demand charges, regardless of interruption risk.
AnswerA

Right-sizing involves analyzing historical utilization metrics (e.g., CloudWatch CPU, memory) to select a less expensive instance type while still satisfying workload requirements. Changing the Auto Scaling group's launch template to a smaller instance type, then monitoring metrics such as CPU credit balance, latency, and throughput, confirms the new size can handle peak demand. This directly lowers per-instance cost without altering the number of instances needed, providing cost savings while preserving availability and performance characteristics.

Why this answer

The application is over-provisioned: CPU utilization averages only 8% with no memory pressure and stable response times. By selecting a smaller EC2 instance type in the Auto Scaling launch template or configuration, you directly reduce the per-instance cost while maintaining adequate performance. This is the most straightforward cost optimization step without modifying the application code or architecture.

Exam trap

The trap here is that candidates may think increasing capacity (Option B) improves efficiency, but in reality, adding more instances to an already underutilized workload only increases cost without any performance benefit.

How to eliminate wrong answers

Option B is wrong because increasing desired capacity to 2x would add more instances, increasing total cost while utilization per instance would drop even further, making the system less efficient, not more. Option C is wrong because disabling Auto Scaling removes the ability to scale down during low demand, which would lock in higher costs and prevent the group from right-sizing to actual load. Option D is wrong because switching to Spot instances immediately without testing or implementing interruption-handling mechanisms (e.g., graceful shutdown, checkpointing) risks application availability and stability, which is not acceptable when the goal is to lower cost without changing the application.

662
MCQmedium

A video platform uses Amazon Aurora. The workload has many short-lived database connections from Lambda functions, causing connection storms. What should be added?

A.S3 Select
B.An internet gateway
C.A larger Route 53 hosted zone
D.RDS Proxy
AnswerD

RDS Proxy is an AWS-managed connection pooler that sits between your application and Aurora, maintaining a warm pool of existing database connections and reusing them across many client sessions. For a bursty video platform workload, this dramatically reduces the CPU and memory overhead per connection and prevents Aurora from exhausting its maximum connection limit. It also keeps connections stable during Aurora failovers and supports IAM authentication, making it the correct choice for this scenario.

Why this answer

RDS Proxy sits between Lambda functions and the Aurora database, pooling and reusing database connections. This prevents the Lambda functions from overwhelming the database with many short-lived connections, which can cause connection storms and degrade performance. RDS Proxy also reduces the overhead of establishing new connections and improves scalability.

Exam trap

The trap here is that candidates may confuse connection pooling with network-level components (like internet gateways) or data retrieval services (like S3 Select), overlooking that RDS Proxy is the AWS-native solution for managing short-lived database connections from serverless or highly concurrent workloads.

How to eliminate wrong answers

Option A is wrong because S3 Select is used to retrieve subsets of data from objects in Amazon S3 using SQL expressions, not for managing database connections. Option B is wrong because an internet gateway enables VPC resources to communicate with the internet, not to manage or pool database connections. Option C is wrong because a larger Route 53 hosted zone increases the number of DNS records you can create, but does not address connection pooling or database connection storms.

663
MCQeasy

A company runs a stateless web application on Amazon EC2 instances behind an Application Load Balancer. The application stores session state in a relational database, which is becoming a bottleneck during peak hours. The team wants to improve performance and reduce database load while keeping the application stateless. Which solution should a solutions architect recommend?

A.Migrate the session table to Amazon DynamoDB with on-demand capacity and update the application to use DynamoDB for session storage.
B.Store session state in an Amazon ElastiCache for Redis cluster and configure the application to read and write sessions to the cluster.
C.Increase the size of the relational database instance and add read replicas, then direct session reads to the replicas and writes to the primary.
D.Enable sticky sessions on the Application Load Balancer and store session state in the local memory of each EC2 instance.
AnswerB

ElastiCache for Redis provides low-latency, in-memory session storage that scales horizontally and removes the database bottleneck. The application remains stateless because session data is externalized to a shared, highly available cache. This improves response times and reduces relational database load during peak traffic, which is a standard pattern for session management.

Why this answer

Externalizing session state to an in-memory cache such as ElastiCache for Redis removes the relational database bottleneck and keeps the application stateless. Redis provides sub-millisecond latency, high throughput, and optional persistence, making it ideal for session stores. Sticky sessions or database scaling do not achieve the same performance and scalability benefits.

Exam trap

The trap here is assuming that sticky sessions or a larger database solve session state performance, when the real issue is that session data should be moved out of the relational database into a low-latency shared cache.

664
MCQmedium

A production log archive runs continuously on EC2 with predictable usage for the next three years. The team wants a discount while retaining some instance-family flexibility. What should they buy? The design must avoid adding custom operational scripts.

A.S3 Intelligent-Tiering
B.Dedicated Instances
C.Compute Savings Plan
D.Spot Instances only
AnswerC

A Compute Savings Plan commits a specific hourly dollar amount for a one- or three-year term in exchange for up to 72% savings over On-Demand, and it automatically applies to any EC2 instance family, size, OS, or Region, plus Fargate and Lambda. A continuously running production log archive is a steady, predictable workload, so committing to a Compute Savings Plan locks in lower unit costs while preserving the flexibility to resize or change instance types later. This makes it the most suitable option for reducing EC2 spend on an always-on baseline.

Why this answer

The Compute Savings Plan (C) offers the largest discount (up to 66%) in exchange for a commitment to a consistent amount of compute usage (measured in $/hour) for a 1- or 3-year term, while still allowing flexibility across instance families, sizes, OS, tenancy, and regions. This matches the predictable three-year workload and the requirement for instance-family flexibility without custom scripts.

Exam trap

The trap here is that candidates confuse Savings Plans with Reserved Instances, assuming that any commitment requires locking into a specific instance family, but Compute Savings Plans explicitly provide family flexibility while still delivering a significant discount.

How to eliminate wrong answers

Option A is wrong because S3 Intelligent-Tiering is an object storage class for data with changing access patterns, not a compute pricing model for EC2 instances. Option B is wrong because Dedicated Instances provide physical isolation at a higher cost and do not offer a discount or instance-family flexibility; they are for compliance or licensing needs. Option D is wrong because Spot Instances offer deep discounts but can be interrupted with a 2-minute warning, making them unsuitable for a production log archive that must run continuously without disruption.

665
MCQmedium

An application writes to an Amazon Aurora DB cluster. After a planned Aurora failover, the application experiences several minutes of connection errors. The logs show the application continues connecting to the specific DB instance endpoint that was the primary before the failover. What change most directly improves resilience during Aurora failovers?

A.Update the application to use the Aurora cluster writer endpoint for write traffic so it always resolves to the current writer instance.
B.Increase Aurora storage autoscaling so failovers are unnecessary.
C.Point both reads and writes to the Aurora reader endpoint to keep the DNS name the same.
D.Disable Aurora failover capability so the cluster never switches writer instances.
AnswerA

During failover, Aurora changes which underlying DB instance is the writer. The cluster writer endpoint (for the cluster) always resolves to the current writer. Using the writer endpoint prevents the application from being pinned to an old instance endpoint that may stop accepting writes after failover.

Why this answer

The Aurora cluster writer endpoint always resolves to the current primary DB instance, even after a failover. By using this endpoint instead of a specific instance endpoint, the application automatically reconnects to the new writer without manual intervention or connection errors.

Exam trap

The trap here is that candidates may think using any Aurora endpoint (like the reader endpoint) is sufficient, but they must understand that only the cluster writer endpoint guarantees write availability after a failover, while the reader endpoint is strictly for read traffic.

Why the other options are wrong

B

Increasing Aurora storage autoscaling does not prevent failovers; it only adjusts storage capacity. Failovers occur due to instance-level issues, not storage limits, so this change does not address the application's connection errors after failover.

C

The reader endpoint is intended for read-only traffic and does not handle write operations; pointing writes to it would cause failures. Moreover, the reader endpoint resolves to multiple reader instances, not the current writer, so it does not ensure connectivity to the writer after failover.

D

Disabling failover capability prevents the cluster from switching to a healthy instance during a failure, which would cause prolonged downtime instead of resolving the connection errors. The question asks for improving resilience, and disabling failover directly undermines that goal.

When would these options actually be correct?

B

A question where an application experiences write failures due to storage capacity limits on an Aurora cluster. The correct solution would be to enable storage autoscaling to automatically increase storage when thresholds are reached, preventing write disruptions.

C

In a scenario where the application only performs read operations and needs to distribute load across replicas while maintaining a single DNS name that automatically adjusts to available instances, using the reader endpoint would be correct. For example, a reporting application that queries read replicas and must remain available during failover.

D

In a scenario where the application cannot tolerate any connection interruption and the database must remain on a specific instance for compliance or licensing reasons, disabling failover might be chosen to avoid automatic failover, accepting the risk of manual recovery.

Why candidates pick the wrong answer

B

Candidates may mistakenly believe that storage autoscaling can eliminate the need for failovers by preventing resource exhaustion, but failovers are triggered by instance health, not storage.

C

Candidates may think that using a single endpoint simplifies DNS resolution and avoids the connection errors seen in the question, but they overlook that the reader endpoint is not designed for write traffic and does not point to the writer instance.

D

Candidates may think that disabling failover eliminates the need for the application to reconnect, thus avoiding connection errors, but they overlook that this removes the cluster's high availability and would cause extended downtime if the primary fails.

666
MCQhard

Based on the exhibit, a public API is behind CloudFront and is experiencing bursts of requests from the same client IP, causing upstream saturation. The team wants AWS to automatically block that IP when the request rate becomes excessive while keeping enforcement as close to the client as possible. Which control should they add?

A.Add an AWS WAF rate-based rule to the CloudFront distribution and configure it to block the source IP after the threshold is exceeded.
B.Add a network ACL rule that denies the source IP after five requests are observed.
C.Enable AWS Shield Advanced and create a custom protection group for the single IP address.
D.Place the API behind a security group rule that allows only the current client IP range.
AnswerA

AWS WAF rate-based rules are purpose-built for this use case. They evaluate the HTTP request rate from a source IP over a sliding window and can automatically block, CAPTCHA, or count when the threshold is exceeded. Attaching the Web ACL to CloudFront enforces the control at the edge, so abusive requests are stopped before they reach the origin and consume upstream capacity.

Why this answer

AWS WAF rate-based rules are designed to automatically block IP addresses that exceed a specified request rate within a 5-minute evaluation window. By attaching this rule to a CloudFront distribution, enforcement occurs at the edge location closest to the client, preventing excessive requests from reaching the upstream API and mitigating saturation.

Exam trap

The trap here is confusing stateless network ACLs or static security groups with the automatic, rate-aware blocking capability of AWS WAF, leading candidates to choose a manual or non-scalable solution.

How to eliminate wrong answers

Option B is wrong because network ACLs are stateless and require manual intervention to add or remove rules; they cannot automatically block an IP after a threshold of requests is observed. Option C is wrong because AWS Shield Advanced provides DDoS protection and custom protection groups for resource-level mitigation, not automatic per-IP rate limiting based on request count. Option D is wrong because security group rules are stateful and cannot dynamically update to block a specific client IP based on request rate; they only allow or deny traffic based on static rules.

667
MCQmedium

A retail company runs a stateless web tier on Amazon EC2 instances behind an Application Load Balancer. Traffic is steady during the day but drops to near zero between 01:00 and 06:00, and the team wants to reduce cost without manual intervention. The instances take about four minutes to boot and warm up. Which configuration meets these requirements?

A.Enable an Application Load Balancer idle timeout and let the load balancer stop idle targets automatically.
B.Attach a step scaling policy that removes instances when the ALB request count per target falls below a threshold.
C.Configure a target tracking scaling policy based on average CPU utilization with a target of 50 percent.
D.Create a scheduled scaling action that reduces the Auto Scaling group's desired capacity at 01:00 and increases it at 06:00.
AnswerD

Scheduled scaling changes desired capacity at known times, which matches a workload with a predictable daily trough. Because the team knows exactly when traffic falls and rises, a schedule avoids the lag of reactive policies and removes the need for manual changes, while the warm-up time is handled by starting the increase slightly before 06:00.

Why this answer

Scheduled scaling is the right fit when capacity needs follow a known timetable: it changes desired capacity at 01:00 and 06:00 regardless of metric lag, so the group shrinks during the predictable trough and is already expanding before morning traffic arrives. Reactive policies such as target tracking or step scaling would respond only after metrics move, which conflicts with the fixed daily pattern and the four-minute warm-up.

Exam trap

The trap here is choosing a metric-driven policy for a workload whose demand is already known by clock time, so the group reacts late instead of scaling ahead of the pattern.

668
MCQhard

A media processing workflow generates analytics files that are accessed unpredictably. Some files become hot again months later. The team wants automatic storage cost optimisation without retrieval delays. What should be used? The architecture review board prefers a managed AWS-native control.

A.S3 Intelligent-Tiering
B.Manual monthly review and object copying
C.S3 Glacier Flexible Retrieval for all files
D.EFS One Zone for analytics files
AnswerA

S3 Intelligent-Tiering is the optimal choice because it automatically monitors access patterns per object and moves data to lower-cost tiers (Infrequent Access, Archive Instant Retrieval) while keeping retrieval latency in milliseconds. It charges a small monthly monitoring fee but avoids retrieval fees and does not require lifecycle rules or retrieval delays, making it ideal for analytics files that may be accessed hot, then cool, and occasionally need immediate access.

Why this answer

S3 Intelligent-Tiering is the correct choice because it automatically moves objects between access tiers (frequent, infrequent, and archive instant retrieval) based on changing access patterns, without any retrieval delays for hot objects. This matches the unpredictable access pattern where files may become hot again months later, and it is a fully managed AWS-native solution that optimizes storage costs automatically.

Exam trap

The trap here is that candidates may choose S3 Glacier Flexible Retrieval (Option C) thinking it is the cheapest for all files, but they overlook the retrieval delay requirement and the fact that files may become hot again, which Intelligent-Tiering handles seamlessly without any retrieval latency.

How to eliminate wrong answers

Option B is wrong because manual monthly review and object copying is not automated, introduces operational overhead, and risks human error or delays, failing the 'automatic' and 'managed AWS-native' requirements. Option C is wrong because S3 Glacier Flexible Retrieval has retrieval delays (minutes to hours) for all files, which violates the 'no retrieval delays' requirement for files that become hot again. Option D is wrong because EFS One Zone is a file system, not an object storage service, and it is not designed for cost optimization of unpredictable access patterns; it also lacks the automatic tiering capability and is not the right service for analytics files that are accessed via S3 APIs.

669
MCQeasy

A company stores 500 TB of archival data in Amazon S3. The data is accessed only for compliance audits, which occur once every two years. Retrieval times of up to 12 hours are acceptable. The company wants the lowest storage cost. Which S3 storage class should be used?

A.S3 One Zone-Infrequent Access (S3 One Zone-IA)
B.S3 Intelligent-Tiering
C.S3 Glacier Deep Archive
D.S3 Standard-Infrequent Access (S3 Standard-IA)
AnswerC

S3 Glacier Deep Archive is the lowest-cost storage class for long-term retention, designed for data accessed less than once per year. It supports retrieval within 12 hours, meeting the acceptable retrieval time. For compliance data accessed only every two years, it provides the most cost-effective solution.

Why this answer

S3 Glacier Deep Archive is specifically designed for long-term retention of data that is rarely accessed, offering the lowest storage cost among S3 classes. It supports retrieval within 12 hours, which matches the acceptable retrieval time. For compliance data accessed every two years, it provides the most economical solution without unnecessary retrieval speed or monitoring overhead.

Exam trap

The trap here is choosing S3 Intelligent-Tiering for automatic cost optimization, but it adds monitoring fees and may not be the absolute lowest cost for data that is predictably cold.

670
MCQmedium

A media company runs a 24/7 recommendation engine on EC2 in one AWS Region. The workload is interruption-intolerant, and the team expects steady usage but may change instance families and sizes during planned optimizations. Compared to the current On-Demand setup, they want the lowest cost while avoiding the rigidity of locking to a specific instance type. What should the solutions architect recommend?

A.Switch the instances to Spot Instances and use interruption handling because it is the largest discount.
B.Purchase a Compute Savings Plan for the expected steady hourly usage in that Region.
C.Purchase a Standard Reserved Instance tied to a single specific instance type for the next 3 years.
D.Keep On-Demand and rely on Auto Scaling to reduce capacity when utilization is low.
AnswerB

Compute Savings Plans offer lower hourly rates in exchange for a 1- or 3-year commitment, but the discount applies to any EC2 instance family/type/size in the chosen Region (even Fargate/Lambda). Because the recommendation engine runs 24/7, committing to the expected steady hourly usage captures the discount while preserving the ability to change instance families or sizes as needs evolve. Unlike Standard RIs, you're not locked into a specific instance type, so you get both cost savings and operational flexibility.

Why this answer

B is correct because a Compute Savings Plan offers the lowest cost for steady-state workloads without locking to a specific instance type, providing up to 66% discount compared to On-Demand while allowing flexibility to change instance families, sizes, OS, or tenancy within a Region. This matches the requirement for cost savings and flexibility during planned optimizations.

Exam trap

The trap here is that candidates often choose Spot Instances for cost savings without considering the interruption-intolerant requirement, or they select Standard Reserved Instances for the highest discount without recognizing the rigidity penalty for planned instance family changes.

Why the other options are wrong

A

The workload is interruption-intolerant, so Spot Instances are unsuitable because they can be terminated with little notice, risking service disruption.

C

A Standard Reserved Instance locks to a specific instance type, which conflicts with the requirement to change instance families and sizes during planned optimizations.

D

The workload is steady and interruption-intolerant, so Auto Scaling to reduce capacity when utilization is low would not provide the lowest cost for the steady baseline usage, and On-Demand pricing is more expensive than a Compute Savings Plan for predictable workloads.

When would these options actually be correct?

A

A question where the workload is fault-tolerant, can handle interruptions (e.g., batch processing, data analysis), and cost reduction is the top priority, with no requirement for steady, uninterruptible compute.

C

A company has a predictable, steady-state workload that uses a specific instance type for 1-3 years, with no plans to change instance families or sizes, and seeks the maximum discount over On-Demand.

D

A company has a variable workload that experiences predictable low-usage periods (e.g., nightly or seasonal drops) and can tolerate scaling down capacity during those times. The goal is to minimize costs by paying only for what is used, without upfront commitments.

Why candidates pick the wrong answer

A

Candidates see 'largest discount' and assume Spot is always best for cost savings, overlooking the critical requirement of interruption intolerance.

C

Candidates may assume Reserved Instances always offer the best savings without considering the flexibility constraint, or they overlook the requirement to change instance types.

D

Candidates may think Auto Scaling always reduces costs by matching capacity to demand, but they overlook that On-Demand pricing is still higher than savings plans for steady usage, and the question explicitly seeks the lowest cost.

671
MCQeasy

A website serves versioned JavaScript and CSS files through CloudFront, but origin fetches are still high and the CloudFront bill increased. Developers confirm that URLs include a version in the filename (for example, app.1.4.2.js). What CloudFront behavior/configuration is most likely to reduce origin fetches and associated costs?

A.Set long cache headers (for example, Cache-Control: max-age and immutable) on those versioned assets so CloudFront caches them longer.
B.Disable compression to reduce CPU time spent at the edge and therefore reduce total cost.
C.Lower the cache policy TTLs so clients always get the newest assets quickly.
D.Remove version identifiers from filenames so CloudFront caches fewer unique objects.
AnswerA

Versioned filenames (e.g., app-1a2b3c.js) enable CloudFront to treat each URL as immutable. Setting Cache-Control: max-age=31536000, immutable on the origin tells CloudFront and browsers to cache the object for a full year without revalidation. Because every new release uses a new URL, stale content is never served, and future requests hit CloudFront's edge cache instead of going to the origin, improving cache hit ratio and reducing transfer cost.

Why this answer

Setting long cache headers like `Cache-Control: max-age=31536000, immutable` on versioned assets tells CloudFront to cache these objects at edge locations for an extended period. Since the filename changes with each new version, CloudFront treats each version as a unique object and will not re-fetch the old version from the origin, dramatically reducing origin fetches and associated costs.

Exam trap

The trap here is that candidates may think lowering TTLs or removing versioning helps with freshness or cost, but the key insight is that versioned filenames already solve cache invalidation, so extending cache duration is the cost-optimized approach.

How to eliminate wrong answers

Option B is wrong because disabling compression does not reduce CPU time at the edge in a meaningful way for cost reduction; CloudFront charges for data transfer and requests, not CPU, and compression actually reduces data transfer costs. Option C is wrong because lowering cache policy TTLs would cause CloudFront to re-fetch objects from the origin more frequently, increasing origin fetches and costs, which is the opposite of the desired outcome. Option D is wrong because removing version identifiers would cause CloudFront to treat all updates as the same object, leading to cache invalidation issues and potentially higher origin fetches when clients request the latest version without a cache busting mechanism.

672
MCQeasy

A company runs an Amazon RDS for PostgreSQL database. The application performs frequent OLTP writes, but it also has a separate dashboard that runs heavy SELECT queries and is slowing down overall database performance. The writes must remain on the primary. What is the best approach to improve performance for the dashboard?

A.Create an RDS read replica and route the dashboard’s read-only queries to the replica endpoint
B.Increase instance storage throughput limits and disable synchronous replication to speed up all queries
C.Replace RDS with Amazon S3 because dashboards require SQL result caching
D.Move the primary database to a different AWS Region to reduce network latency
AnswerA

Creating an RDS read replica gives the dashboard an independent database endpoint that serves read-only SELECT traffic without consuming the primary instance's CPU, memory, or I/O capacity. PostgreSQL replication is asynchronous, so replica data is eventually consistent, but that is generally acceptable for dashboards that tolerate a few seconds of lag. By routing dashboard connections to the replica's DNS endpoint, OLTP write transactions keep using the primary, reducing lock and buffer-pool contention and improving overall application responsiveness.

Why this answer

Creating an RDS read replica allows you to offload the heavy SELECT queries from the primary database instance. The replica asynchronously replicates data from the primary using PostgreSQL's streaming replication, so the dashboard can query the replica without impacting the OLTP write performance on the primary. This directly addresses the requirement that writes remain on the primary while improving dashboard query performance.

Exam trap

The trap here is that candidates might think increasing instance size or storage throughput is sufficient, but the core issue is workload isolation—offloading read-heavy queries to a read replica is the only scalable solution that preserves write performance on the primary.

Why the other options are wrong

B

Disabling synchronous replication compromises data durability and is not supported for RDS PostgreSQL; increasing storage throughput does not address the root cause of heavy SELECT queries impacting OLTP writes.

C

Amazon S3 is an object storage service, not a relational database, and cannot execute SQL queries or replace the transactional and query capabilities of RDS PostgreSQL. The dashboard requires live querying of the same data, not cached results from S3.

D

Moving the primary database to a different AWS Region does not address the performance impact of heavy SELECT queries on the same instance; it only changes the geographic location, potentially increasing latency for writes and reads.

When would these options actually be correct?

B

An exam scenario where the database is experiencing I/O bottlenecks due to insufficient throughput for both reads and writes, and the question allows modifying replication settings to prioritize write performance at the cost of durability.

C

A question where the requirement is to store and serve large amounts of static or semi-static data for a dashboard that can tolerate eventual consistency, and the dashboard queries are simple key-value lookups or use Athena/Redshift Spectrum for analytics on data stored in S3.

D

This option would be correct if the question stated that the primary database is in a region far from the application servers, causing high latency for all operations, and the goal is to reduce network latency by relocating the database closer to the application.

Why candidates pick the wrong answer

B

Candidates may think that increasing resources or reducing replication overhead can solve performance issues, without understanding that read replicas are the proper solution for offloading read traffic without affecting writes.

C

Candidates may think S3 is a cost-effective solution for dashboards because it can store large datasets and integrate with analytics services, but they overlook that S3 is not a database and cannot handle complex SQL queries or transactional workloads.

D

Candidates may think that moving to a different region can reduce latency for the dashboard, but they overlook that the primary database still handles all queries, and the dashboard's heavy SELECTs remain on the same instance, not solving the performance issue.

673
MCQhard

Based on the exhibit, an application runs in private subnets without a NAT gateway and must retrieve a secret from AWS Secrets Manager. Security requires the traffic to stay on the AWS network and not traverse the public internet. What is the best solution?

A.Add a NAT gateway to the private subnet route table and keep using the public Secrets Manager endpoint.
B.Create an interface VPC endpoint for Secrets Manager and enable private DNS for the endpoint.
C.Create a gateway VPC endpoint for Secrets Manager and point the route table to it.
D.Use VPC peering to connect the application subnet to another VPC that already has internet access.
AnswerB

An interface VPC endpoint for Secrets Manager uses AWS PrivateLink to place an elastic network interface with a private IP directly into your subnet, making the service reachable without internet access. Enabling private DNS for the endpoint ensures the default Secrets Manager DNS name resolves to that private interface instead of the public endpoint, so your application code works unchanged. This keeps traffic entirely within the AWS network, avoids internet transit, and requires no NAT or internet gateway.

Why this answer

An interface VPC endpoint for Secrets Manager allows the application in the private subnet to securely access Secrets Manager over the AWS network using private IP addresses, without needing a NAT gateway or internet gateway. Enabling private DNS ensures that the default Secrets Manager DNS name resolves to the endpoint's private IP addresses, keeping all traffic within the AWS backbone and satisfying the security requirement.

Exam trap

The trap here is that candidates confuse gateway VPC endpoints (which work only for S3 and DynamoDB) with interface VPC endpoints (which are used for most other AWS services including Secrets Manager), leading them to incorrectly select option C.

How to eliminate wrong answers

Option A is wrong because adding a NAT gateway would route traffic to the public Secrets Manager endpoint over the internet, violating the requirement that traffic must not traverse the public internet. Option C is wrong because gateway VPC endpoints are only supported for AWS services like S3 and DynamoDB, not for Secrets Manager, which requires an interface endpoint. Option D is wrong because VPC peering with another VPC that has internet access still requires the application to go through a NAT or internet gateway to reach Secrets Manager, breaking the 'no public internet' rule and adding unnecessary complexity.

674
MCQmedium

A company uses Amazon RDS with automated backups enabled (retention period: 7 days). At 10:30 UTC, a bad release corrupts specific rows in a production table. The team detects the issue at 11:10 UTC. They need to revert the database state to what it was from 10:00–10:30 UTC, recover quickly, and minimize risk to the currently running workload. What is the best option?

A.Reboot the DB instance and rely on the corrupted data being overwritten by storage-level changes.
B.Perform a point-in-time restore to a new DB instance using a timestamp before the corruption (for example, a time within 10:00–10:30 UTC).
C.Restore only the most recent automated backup snapshot, even if it is after the corruption timestamp.
D.Create a read replica of the current DB instance and overwrite the corrupted table using SELECT queries from the replica.
AnswerB

With automated backups enabled, RDS supports point-in-time recovery (PITR) within the retention window. Restoring to a timestamp before the corruption creates a consistent copy from that moment. The team can validate the restored DB and then cut over application traffic, reducing risk to the currently running workload.

Why this answer

Amazon RDS automated backups enable point-in-time recovery (PITR) to any second within the retention window. By restoring to a timestamp between 10:00 and 10:30 UTC, you recover the database to a state before the corruption occurred, without affecting the current production instance. This minimizes risk to the running workload because the restore creates a new DB instance, leaving the original untouched until you are ready to switch.

Exam trap

The trap here is that candidates may confuse automated backup snapshots (which are full backups taken once per day) with point-in-time recovery (which uses transaction logs to restore to any point within the retention window), leading them to choose Option C instead of B.

How to eliminate wrong answers

Option A is wrong because rebooting a DB instance does not revert data; it only restarts the database engine and does not undo committed transactions or storage-level changes. Option C is wrong because restoring the most recent automated backup snapshot includes the corrupted data, so it does not achieve the goal of reverting to a pre-corruption state. Option D is wrong because a read replica mirrors the current (corrupted) data; using SELECT queries from it cannot overwrite the corrupted table with clean data, and it does not provide a mechanism to roll back changes.

675
MCQmedium

A service processes customer payments from a message queue. Because the queue provides at-least-once delivery, the same payment message can be delivered more than once if the consumer times out before committing its state. Currently, the service sometimes charges the customer twice. Which design change most directly prevents duplicate charges while still allowing safe retries?

A.Delete the message from the queue immediately after receive to prevent redelivery.
B.Make the payment processing idempotent by recording an idempotency key for each payment and ensuring repeated deliveries do not apply the charge twice.
C.Increase the queue visibility timeout to a very large value so messages rarely reappear.
D.Switch to a single-threaded consumer with one worker so messages are processed in order.
AnswerB

Because message queues provide at-least-once delivery, a payment message may be delivered multiple times if a consumer times out or fails after processing. By writing the payment's idempotency key (for example, a payment reference) to a durable store with a unique constraint before applying the charge, the consumer can detect and ignore repeated deliveries. This ensures the charge happens exactly once even when a redelivery occurs, making the system safe against at-least-once semantics.

Why this answer

Making payment processing idempotent using an idempotency key ensures that even if the same message is delivered multiple times due to at-least-once delivery semantics, the charge is applied only once. The consumer records a unique key (e.g., payment ID) in a durable store (like DynamoDB or Redis) and checks it before processing; if the key already exists, the charge is skipped. This directly prevents duplicate charges while still allowing safe retries, as the consumer can safely reprocess messages without side effects.

Exam trap

The trap here is that candidates often confuse at-least-once delivery with exactly-once delivery and assume that increasing visibility timeouts or using single-threaded consumers will prevent duplicates, when in fact only idempotency guarantees safe retries without duplicate charges.

Why the other options are wrong

A

Deleting the message immediately after receive prevents redelivery but also eliminates the ability to retry if processing fails, which violates the requirement of allowing safe retries.

C

Increasing the visibility timeout to a very large value does not guarantee that a consumer won't crash or timeout, and it can delay processing of other messages, leading to potential bottlenecks and still allowing duplicate charges if the consumer fails after processing but before deleting the message.

D

Single-threading does not prevent duplicate charges because the same message can still be redelivered after a timeout, even with one worker; the core issue is at-least-once delivery, not concurrency.

When would these options actually be correct?

A

If the question required exactly-once processing with no retries and the consumer could guarantee successful processing upon receipt, deleting immediately would be correct.

C

This option would be correct in a scenario where the queue is used for time-sensitive tasks that must not be processed concurrently, and the consumer is highly reliable with minimal failure risk, such as processing a single critical job that should not be retried quickly.

D

This option would be correct in a scenario where the queue guarantees exactly-once delivery but the consumer needs to maintain strict ordering and avoid race conditions from multiple concurrent workers.

Why candidates pick the wrong answer

A

Candidates may think that eliminating redelivery directly solves duplicate charges, overlooking the need for retries in case of failures.

C

Candidates may think that a larger visibility timeout prevents message redelivery during normal processing, overlooking that failures or timeouts can still occur, and that this approach does not address the root cause of duplicate processing due to at-least-once delivery semantics.

D

Candidates may think that processing messages sequentially eliminates duplicates, but they overlook that at-least-once delivery can still cause redelivery of the same message to a single-threaded consumer.

Page 8

Page 9 of 13

Page 10