Courseiva

SAA-C03 (SAA-C03) — Questions 901–935

935 questions total · 13pages · All types, answers revealed

Page 12

Page 13 of 13

901
MCQmedium

A company runs an internal analytics application in a single AWS Region. A solutions architect is reviewing the Amazon RDS for MySQL deployment and finds a Multi-AZ DB instance with a standby in another Availability Zone, used only for failover. The application performs many read-heavy queries against the primary instance, driving up instance size and cost. The team wants to offload read traffic and reduce the primary instance size. Which change should the architect recommend?

A.Convert the deployment to a Multi-AZ DB cluster with two readable standbys.
B.Migrate the database to Amazon DynamoDB with on-demand capacity.
C.Add RDS read replicas and direct read-heavy queries to them.
D.Enable RDS Performance Insights and rely on it to reduce query cost.
AnswerC

RDS read replicas serve read-only traffic on separate instances, letting the team offload read-heavy queries and shrink the primary instance. This directly addresses the cost driver, which is an oversized primary handling both reads and writes, and it works within a single Region while the Multi-AZ standby continues to provide high availability for failover.

Why this answer

Adding RDS read replicas moves read-heavy queries onto separate instances, relieving the primary and enabling a smaller, cheaper primary instance. The Multi-AZ standby exists only for failover and cannot serve reads, so the architect should introduce read replicas rather than re-architecting the deployment or switching database engines.

Exam trap

The trap here is assuming the Multi-AZ standby can serve read traffic, when a Multi-AZ DB instance standby is strictly for failover and never handles application reads.

902
MCQhard

Based on the exhibit, an application repeatedly reads the same DynamoDB items with extremely low latency requirements. The business can tolerate data that is a few seconds stale. Which architecture change best improves read performance?

A.Add a DynamoDB Accelerator (DAX) cluster in front of the table.
B.Increase the table's sort key cardinality while keeping the same read pattern.
C.Switch the table to provisioned mode with auto scaling disabled.
D.Move the session data to Amazon EFS so the application can read it from shared files.
AnswerA

DAX is an in-memory cache designed specifically for DynamoDB, and it can serve repeated reads of the same item with sub-millisecond latency without consuming read capacity units. Because the application's workload is dominated by repeatedly reading the same session data, DAX will yield a high cache hit rate, offload the underlying table, and reduce throttling under peak demand. This makes it the most direct and lowest-application-change solution.

Why this answer

Adding a DynamoDB Accelerator (DAX) cluster provides an in-memory cache that can reduce read latencies to microseconds for frequently accessed items, while still allowing for eventual consistency and tolerating a few seconds of staleness. DAX is specifically designed for this use case, handling cache hits without any application code changes and offloading read traffic from the DynamoDB table.

Exam trap

The trap here is that candidates may overlook DAX as a specialized caching layer for DynamoDB and instead consider increasing table capacity or changing data models, which do not directly address the need for extremely low latency on repeated reads of the same items.

How to eliminate wrong answers

Option B is wrong because increasing sort key cardinality does not improve read performance for repeated reads of the same items; it primarily helps with write distribution and query flexibility, not latency for individual GetItem operations. Option C is wrong because switching to provisioned mode with auto scaling disabled does not inherently improve read performance; it may lead to throttling if capacity is insufficient, and it does not address the need for sub-millisecond latency. Option D is wrong because moving session data to Amazon EFS introduces file system overhead and network latency that is significantly higher than DynamoDB's single-digit millisecond latency, and EFS is not designed for the same low-latency, high-throughput access pattern required for repeated reads of individual items.

903
MCQeasy

An internal API is hosted in two AWS Regions behind Route 53. Under normal conditions, clients should use the primary region. If the primary endpoint becomes unhealthy, traffic must automatically switch to the secondary region. Which Route 53 setup best meets this requirement?

A.Use latency-based routing with one record per region and no health checks.
B.Use failover routing policy: create two alias records for the same name (primary and failover) and associate health checks with the primary record.
C.Use weighted routing and manually change the weights during incidents.
D.Create a single alias record only for the primary region and rely on client-side DNS retries.
AnswerB

Failover routing with two alias records for the same DNS name (one marked primary, one marked secondary) gives you deterministic active-passive failover. You attach a health check to the primary alias; when that health check fails, Route 53 automatically returns the secondary record's endpoint in the next DNS response, without manual intervention. Alias records allow you to point directly to regional load balancers or other AWS resources, and the secondary record ensures that all traffic moves to the healthy region once the primary is considered unhealthy.

Why this answer

Route 53 failover routing policy is designed for active-passive failover scenarios. By creating two alias records (primary and secondary) for the same DNS name and associating a health check with the primary record, Route 53 automatically directs traffic to the secondary region if the primary health check fails. This meets the requirement of automatic failover without manual intervention.

Exam trap

The trap here is that candidates often confuse failover routing with latency-based routing, assuming latency routing inherently handles failover, but latency routing does not automatically switch traffic when an endpoint becomes unhealthy unless health checks are explicitly configured.

How to eliminate wrong answers

Option A is wrong because latency-based routing distributes traffic based on lowest latency, not active-passive failover, and without health checks it cannot detect endpoint failures. Option C is wrong because weighted routing requires manual weight changes during incidents, which violates the requirement for automatic failover. Option D is wrong because a single alias record with no secondary endpoint provides no failover capability; client-side DNS retries do not redirect to a different region.

904
MCQmedium

A Lambda function for a mobile banking backend needs to read a database password. The password must rotate automatically every 30 days and should not be stored in environment variables. Which service should be used? The design must avoid adding custom operational scripts.

A.An encrypted object in Amazon S3
B.AWS Secrets Manager with rotation enabled
C.AWS Systems Manager Parameter Store SecureString without automation
D.A KMS-encrypted Lambda environment variable
AnswerB

Secrets Manager natively rotates credentials on a schedule using built-in Lambda rotation functions, satisfying the 30-day rotation requirement without custom scripts. It also keeps the password out of environment variables, unlike plain Lambda configuration or SSM Parameter Store without rotation logic.

Why this answer

AWS Secrets Manager is the correct choice because it natively supports automatic rotation of secrets on a configurable schedule (e.g., every 30 days) without requiring custom scripts. It also provides fine-grained access control and integrates directly with Lambda via the AWS SDK, keeping the password out of environment variables and code.

Exam trap

The trap here is that candidates often confuse AWS Systems Manager Parameter Store SecureString (which can store secrets but lacks automatic rotation) with Secrets Manager, or they assume that encrypting environment variables with KMS is sufficient for rotation, ignoring the need for automated lifecycle management.

How to eliminate wrong answers

Option A is wrong because storing an encrypted object in Amazon S3 requires custom code to retrieve, decrypt, and rotate the password, violating the 'no custom operational scripts' constraint. Option C is wrong because AWS Systems Manager Parameter Store SecureString without automation does not support automatic rotation; you would need to manually update the parameter or add a custom rotation solution. Option D is wrong because a KMS-encrypted Lambda environment variable is static and cannot be rotated automatically; you would need to redeploy the function to change the password, which adds operational overhead.

905
MCQhard

Based on the exhibit, what change should the team make to achieve the lowest possible network latency for the distributed workload?

A.Place the instances in a spread placement group across multiple Availability Zones.
B.Move the workload into a cluster placement group in one Availability Zone.
C.Add an Application Load Balancer in front of the workers to reduce inter-node latency.
D.Increase the EC2 instance size while keeping the current multi-AZ layout.
AnswerB

Cluster placement groups place instances physically close together inside one Availability Zone, which is the best AWS option for workloads that need low-latency, high-bandwidth communication between many nodes. The exhibit explicitly says the workload can run in a single AZ if performance improves. That makes cluster placement groups the right fit.

Why this answer

A cluster placement group provides the lowest possible network latency and highest throughput by placing all instances in a single Availability Zone with low-latency, non-blocking 10 Gbps or 25 Gbps network connectivity between them. This is ideal for tightly coupled, distributed workloads that require frequent inter-node communication, such as HPC or data analytics jobs.

Exam trap

The trap here is that candidates often assume multi-AZ is always better for high availability, but for latency-sensitive distributed workloads, a single-AZ cluster placement group is the correct choice to minimize inter-node latency, even though it sacrifices fault tolerance.

Why the other options are wrong

A

Spread placement groups maximize availability and fault isolation by placing instances on distinct hardware across multiple AZs, which increases network latency due to physical separation. The goal is lowest possible latency, which requires a cluster placement group in a single AZ.

C

An Application Load Balancer distributes incoming traffic across targets, but it does not reduce inter-node latency between workers; in fact, it adds a hop and increases latency for node-to-node communication.

D

Increasing EC2 instance size does not reduce network latency between instances; it improves compute capacity. The question specifically asks for lowest network latency, which requires physical proximity, not larger instances.

When would these options actually be correct?

A

A question asks for the highest availability and fault tolerance for a critical application that must survive an entire AZ failure, and latency is not the primary concern. Spread placement groups across AZs would be correct.

C

This option would be correct in a question asking how to distribute incoming client requests across multiple EC2 instances in different Availability Zones for high availability and fault tolerance, while also providing health checks and SSL termination.

D

This option would be correct in a scenario where the workload is compute-bound and requires more CPU or memory per instance, and the goal is to improve throughput or reduce processing time, not network latency.

Why candidates pick the wrong answer

A

Candidates may think that distributing instances across AZs always improves performance, confusing high availability with low latency, or they may overvalue fault tolerance when the question prioritizes latency.

C

Candidates may think that load balancers always improve performance, confusing load distribution for user-facing traffic with reducing latency for internal distributed workloads.

D

Candidates may assume larger instances have better network performance or that more resources inherently reduce latency, confusing compute capacity with network characteristics.

906
MCQeasy

A startup runs a public-facing web application on Amazon EC2 instances behind an Application Load Balancer. The application must call AWS APIs such as Amazon DynamoDB and Amazon S3. A security engineer must ensure that no long-term AWS credentials are stored on the instances and that each instance receives credentials automatically. Which solution should the engineer use?

A.Store an IAM user access key and secret key in AWS Systems Manager Parameter Store as a SecureString and retrieve them at instance boot.
B.Configure the instances to call AWS Security Token Service AssumeRole with a shared secret stored in an encrypted Amazon EBS volume.
C.Create an IAM user per instance, generate access keys, and rotate them every 90 days using a scheduled AWS Lambda function.
D.Attach an IAM role for Amazon EC2 to an instance profile and associate it with each instance so the SDK retrieves temporary credentials from the instance metadata service.
AnswerD

An instance profile delivers temporary credentials through the instance metadata service, and the AWS SDK refreshes them automatically before they expire. No access keys are written to disk or configuration files, and permissions are controlled by the role's policies, so the design eliminates long-term credentials while granting least-privilege access to DynamoDB and S3.

Why this answer

Associating an IAM role with the instances through an instance profile lets the AWS SDK obtain temporary, automatically rotated credentials from the instance metadata service. This removes the need to store access keys on the instances, limits permissions to what the role allows, and satisfies the requirement that each instance receives credentials without manual handling.

Exam trap

The trap here is treating an encrypted store such as Parameter Store or an encrypted EBS volume as sufficient, when the underlying problem is that the credentials themselves are long-lived rather than the storage medium.

907
MCQmedium

A company stores millions of objects in Amazon S3. Access patterns are completely unpredictable — some objects are frequently accessed, others rarely. Objects range from 4 KB to 50 MB. The company wants to minimize storage costs automatically without managing lifecycle rules. Which storage class should a solutions architect recommend?

A.S3 Standard — it is the default and handles all access patterns equally
B.S3 Standard-IA — it automatically detects infrequent access and reduces cost
C.S3 Intelligent-Tiering — it automatically moves objects between tiers based on access patterns
D.S3 One Zone-IA — it is the cheapest option with fast retrieval
AnswerC

S3 Intelligent-Tiering continuously monitors access at the object level and automatically shifts objects between frequent, infrequent, and optional archive-access tiers based on recent usage, all without retrieval fees and without requiring lifecycle transitions. It charges only a small monthly automation/monitoring fee per object, which is economical when access behavior is unknown. This idle-tier management matches unpredictable workloads precisely, unlike the fixed tiering of Standard-IA or the constant high cost of Standard.

Why this answer

S3 Intelligent-Tiering monitors access patterns and automatically moves objects between access tiers — Frequent Access, Infrequent Access, and optional Archive tiers — based on actual usage. It requires no management or lifecycle rules.

Important: Intelligent-Tiering charges a small monitoring fee per object per month. For objects under 128 KB, this fee may exceed the storage savings. With objects ranging from 4 KB to 50 MB and unpredictable access patterns, Intelligent-Tiering is the recommended answer — AWS explicitly recommends it for unknown access patterns where object size averages above 128 KB.

Exam trap

For purely small objects (all < 128 KB), Intelligent-Tiering's monitoring cost ($0.0025 per 1,000 objects) can exceed the storage savings — Standard would be cheaper. But for mixed sizes with unpredictable access (as in this question), Intelligent-Tiering is the correct recommendation. The key phrase 'automatically without managing lifecycle rules' points to Intelligent-Tiering.

Why the other options are wrong

A

S3 Standard is the highest cost per-GB storage class and does not automatically reduce cost based on access patterns. For unpredictable access, Intelligent-Tiering is more cost-effective for objects with average size above 128 KB.

B

S3 Standard-IA does NOT automatically detect access patterns. Objects placed in Standard-IA are statically in that class. It also charges a per-GB retrieval fee making it expensive for frequently accessed objects.

D

One Zone-IA stores data in a single AZ (lower durability). It does not automatically adjust to access patterns and charges retrieval fees. It's inappropriate for data requiring standard S3 durability.

908
MCQmedium

A read-heavy document portal repeatedly queries the same product catalogue data from DynamoDB with millisecond latency requirements. Which service can reduce read latency and table load? The design must avoid adding custom operational scripts.

A.Amazon Kinesis Data Firehose
B.S3 Transfer Acceleration
C.DynamoDB Accelerator (DAX)
D.AWS Glue Data Catalog
AnswerC

DynamoDB Accelerator (DAX) is a fully managed, highly available in-memory cache that sits in front of DynamoDB and transparently intercepts GetItem, BatchGetItem, and Query calls, returning cached items in microseconds. It uses a write-through strategy that keeps the cache coherent with the underlying table, making it ideal for read-heavy workloads that repeatedly query the same items, like this document portal. DAX also reduces read capacity unit consumption and protects the table from unexpected throttling on hot keys.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache for DynamoDB that delivers microsecond read latency, reducing the number of read requests hitting the underlying table. It requires no custom scripts—just a DAX cluster endpoint—and automatically caches frequently accessed items, making it ideal for a read-heavy document portal with millisecond latency requirements.

Exam trap

The trap here is that candidates often confuse caching services like ElastiCache with DAX, but DAX is purpose-built for DynamoDB and requires no application code changes beyond pointing to a different endpoint, whereas ElastiCache would need custom cache invalidation logic.

How to eliminate wrong answers

Option A is wrong because Amazon Kinesis Data Firehose is a streaming data ingestion service for loading data into data lakes or analytics tools, not a read cache for DynamoDB. Option B is wrong because S3 Transfer Acceleration speeds up uploads to S3 over long distances using edge locations, but does not reduce read latency or load on a DynamoDB table. Option D is wrong because AWS Glue Data Catalog is a metadata repository for ETL jobs and data discovery, not a caching layer for DynamoDB reads.

909
MCQeasy

A startup runs a public-facing web application on Amazon EC2 instances in a VPC. The security team wants to protect the application from common web exploits such as SQL injection and cross-site scripting, and also wants to block traffic from specific countries. Which AWS service should a solutions architect use?

A.AWS WAF with a web ACL attached to the Application Load Balancer, using managed rule groups and a geo match rule.
B.A network ACL on the public subnet that denies traffic from specific country IP ranges.
C.Amazon GuardDuty with S3 protection and a threat intelligence feed.
D.AWS Shield Advanced with a web ACL attached to the Application Load Balancer.
AnswerA

AWS WAF inspects HTTP requests and can block SQL injection and cross-site scripting using AWS managed rule groups. It also supports a geo match condition to block requests from specified countries, directly satisfying both requirements when attached to the ALB.

Why this answer

AWS WAF is the service that inspects HTTP requests for web exploits and can apply geo match conditions. Attaching a web ACL to the Application Load Balancer lets WAF evaluate each request against managed rule groups for SQL injection and cross-site scripting and against a geo match rule that blocks specified countries.

Exam trap

The trap here is confusing network-layer filtering, such as network ACLs, with application-layer inspection that only AWS WAF performs.

910
MCQeasy

An orders service currently sends HTTP requests directly to two downstream services (inventory and shipping). During peak load, inventory slows down, causing the orders service to slow as well. The team wants the orders service to remain responsive even when a downstream service is temporarily slow or restarted. Which design change best achieves this resiliency goal?

A.Keep HTTP calls but add longer client timeouts so orders requests wait for slow downstream responses.
B.Introduce Amazon SQS as a buffer between orders and downstream services, with consumers processing from the queue.
C.Replace the downstream services with AWS Lambda functions that are invoked synchronously by the orders service.
D.Call the downstream services in parallel threads to reduce waiting time during peak load.
AnswerB

SQS decouples the producer (orders service) from the consumers (inventory/shipping processors). The orders service can quickly enqueue work and return to the caller, even if a downstream service is slow or restarted. Messages remain in the queue until consumers can process them, preventing cascading latency/backpressure from propagating to the orders API.

Why this answer

Introducing Amazon SQS as a buffer decouples the orders service from the downstream inventory and shipping services. The orders service can immediately enqueue messages and respond to the client, while downstream consumers process messages at their own pace. This prevents backpressure from a slow or restarting downstream service from blocking the orders service, achieving the desired resiliency.

Exam trap

The trap here is that candidates may think parallelizing calls (Option D) or increasing timeouts (Option A) solves the problem, but they fail to recognize that true resiliency requires decoupling via asynchronous messaging, not just concurrency or tolerance of delays.

How to eliminate wrong answers

Option A is wrong because adding longer client timeouts does not prevent the orders service from being blocked; it only increases the wait time before a timeout occurs, still causing the orders service to slow down during peak load. Option C is wrong because replacing downstream services with synchronously invoked Lambda functions does not decouple the services; the orders service would still block waiting for the Lambda invocation to complete, and Lambda has a 15-minute timeout limit, which does not solve the slowdown issue. Option D is wrong because calling downstream services in parallel threads reduces latency only if both services are responsive; if one service is slow or restarting, the orders service still waits for that slow response, and thread pool exhaustion can occur under peak load, leading to resource contention and slowdown.

911
MCQhard

A Lambda-based retail API has unpredictable traffic spikes and users see latency caused by cold starts. The function must respond consistently during expected campaign windows. What should be configured?

A.A larger deployment package
B.Reserved concurrency only
C.Provisioned concurrency during campaign windows
D.CloudTrail data events
AnswerC

Provisioned concurrency is correct because it keeps a specified number of Lambda execution environments fully initialized and ready to handle requests immediately, thereby eliminating cold-start latency for traffic within the provisioned capacity. By configuring it only during campaign windows, you align the pre-warmed capacity with predictable traffic surges, which both improves response times and controls cost since you are not paying for idle provisioned concurrency outside those windows. This directly addresses the unpredictable spikes by turning them into a scheduled, managed pattern.

Why this answer

Provisioned concurrency initializes a specified number of execution environments in advance, eliminating cold starts during campaign windows. This ensures consistent latency even under unpredictable traffic spikes, as the function is always warm and ready to handle requests immediately.

Exam trap

The trap here is confusing reserved concurrency (which limits scaling but does not prevent cold starts) with provisioned concurrency (which pre-warms environments to eliminate cold starts).

How to eliminate wrong answers

Option A is wrong because a larger deployment package increases cold start time, making latency worse. Option B is wrong because reserved concurrency only guarantees a maximum number of concurrent executions but does not pre-warm environments; cold starts still occur. Option D is wrong because CloudTrail data events record API activity for auditing, not for managing function initialization or latency.

912
MCQeasy

Your team serves static JavaScript and CSS files from an S3 origin through CloudFront. After a release, the CloudFront cache hit ratio dropped because clients keep re-downloading the same assets. What is the best next change to improve caching performance?

A.Update origin responses to include long-lived Cache-Control headers (for example, max-age) so CloudFront can cache objects
B.Switch the S3 bucket to S3 Glacier so objects are not frequently accessed
C.Disable CloudFront compression to reduce CPU usage at the edge
D.Set CloudFront to forward all query strings to the origin to ensure the latest assets are returned
AnswerA

CloudFront will only reuse cached objects when the origin response is cacheable. Adding/adjusting Cache-Control (and related directives such as public and s-maxage where appropriate) to allow long-lived caching enables edge reuse and increases cache hit ratio.

Why this answer

Setting long-lived Cache-Control headers (e.g., max-age=31536000) on static assets tells CloudFront to cache them at edge locations for an extended period. This reduces the number of requests forwarded to the S3 origin, improving the cache hit ratio and preventing clients from re-downloading unchanged assets on every visit.

Exam trap

The trap here is that candidates may think forwarding query strings (Option D) ensures freshness, but it actually fragments the cache and reduces hit ratio, whereas the real solution is to use long-lived Cache-Control headers with versioned filenames to maximize caching.

How to eliminate wrong answers

Option B is wrong because moving the S3 bucket to Glacier would make objects inaccessible for real-time serving, breaking the static asset delivery entirely. Option C is wrong because disabling CloudFront compression does not affect caching behavior; it would only increase bandwidth and latency for clients, not improve cache hit ratio. Option D is wrong because forwarding all query strings to the origin forces CloudFront to treat each unique query string as a separate cache key, fragmenting the cache and reducing hit ratio, which is the opposite of what is needed.

913
MCQmedium

A content publishing system uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The design must avoid adding custom operational scripts.

A.Lambda reserved concurrency set to zero
B.A larger deployment package
C.CloudFront error pages
D.A Lambda dead-letter queue or failure destination
AnswerD

Configuring a Lambda dead-letter queue (SQS or SNS) or an on-failure destination for asynchronous invocations ensures that events that exhaust Lambda's retry policy are captured and stored. Lambda sends the failed event payload to the chosen DLQ or destination, allowing downstream consumers to analyze, quarantine, or reprocess it. This is the standard, built-in mechanism for persisting asynchronous invocation failures.

Why this answer

A Lambda dead-letter queue (DLQ) or failure destination is the correct solution because it captures events that have exhausted all retry attempts from an asynchronous Lambda invocation. This allows failed events to be retained in an Amazon SQS queue or SNS topic for later investigation, without requiring custom operational scripts. The DLQ or failure destination integrates directly with Lambda's built-in retry behavior, ensuring that only events that fail after the configured number of retries are sent to the destination.

Exam trap

The trap here is that candidates may confuse a DLQ with other error-handling mechanisms like CloudFront error pages or reserved concurrency, but only a DLQ or failure destination directly captures failed asynchronous Lambda events without custom code.

How to eliminate wrong answers

Option A is wrong because setting reserved concurrency to zero would prevent the Lambda function from executing at all, which stops all invocations and does not retain failed events. Option B is wrong because a larger deployment package does not affect error handling or event retention; it only increases the function's storage size and cold start time. Option C is wrong because CloudFront error pages are used for HTTP error responses from a web distribution, not for capturing failed Lambda invocations from asynchronous event sources.

914
Multi-Selectmedium

A healthcare company runs a batch ingestion pipeline on Amazon EC2 instances that read messages from an Amazon SQS queue and write results to Amazon DynamoDB. The pipeline must be resilient so that a single instance failure does not stop processing and no messages are lost. Which two architectural changes should a solutions architect make to meet these requirements? (Choose two.)

Select 2 answers
A.Increase the number of EC2 instances manually to five and place them all in the same Availability Zone for low latency.
B.Convert the SQS queue to a FIFO queue with content-based deduplication enabled to guarantee no message loss.
C.Move the ingestion workers into an Auto Scaling group across multiple Availability Zones with a launch template, so failed instances are replaced automatically.
D.Enable DynamoDB Accelerator (DAX) on the target table to cache writes so failed instances do not lose data.
E.Configure the SQS queue with a visibility timeout longer than the maximum processing time and have workers delete a message only after successful processing.
AnswersC, E

An Auto Scaling group spanning multiple Availability Zones maintains desired capacity and replaces unhealthy instances automatically, removing the single-instance failure point. Combined with a queue-based workload, replacement workers resume consuming messages, so processing continues without manual intervention and the pipeline remains resilient to instance loss.

Why this answer

Resilience here has two parts: redundant compute that self-heals, and queue semantics that return unprocessed messages to the queue. A multi-AZ Auto Scaling group replaces failed workers, while a visibility timeout longer than processing time plus delete-after-success ensures a message is reprocessed rather than lost if an instance dies mid-flight.

Exam trap

The trap here is treating caching or FIFO ordering as durability mechanisms, when message safety actually depends on visibility timeout and delete-after-success semantics.

915
MCQhard

A risk simulation workload in private subnets downloads large amounts of data from S3 through a NAT gateway. NAT data processing charges are high. What should the architect use to reduce cost? The design must avoid adding custom operational scripts.

A.A larger NAT gateway
B.Gateway VPC endpoint for Amazon S3
C.S3 Object Lambda
D.AWS Shield Advanced
AnswerB

A gateway VPC endpoint for S3 installs a prefix list route (e.g., com.amazonaws.region.s3) in the VPC route table, causing S3-bound traffic to be sent directly to S3 over the AWS private network instead of through the NAT gateway. Because the data path no longer passes through the NAT gateway, the per-GB NAT data processing charge for those large downloads is eliminated. The endpoint itself is free, highly available, and requires only route table updates plus an optional endpoint policy to control access, making this the correct cost optimization.

Why this answer

A Gateway VPC Endpoint for Amazon S3 allows instances in private subnets to access S3 directly over the AWS network without traversing a NAT gateway, eliminating NAT data processing charges. This is the most cost-effective and operationally simple solution because it requires no custom scripts and no changes to routing beyond adding the endpoint.

Exam trap

The trap here is that candidates often confuse Gateway VPC Endpoints with Interface VPC Endpoints, assuming both incur hourly charges, or mistakenly think a larger NAT gateway is a cost-saving measure when it actually increases costs.

How to eliminate wrong answers

Option A is wrong because a larger NAT gateway would increase, not reduce, data processing costs (charged per GB processed) and does not address the root cause of traffic going through the NAT. Option C is wrong because S3 Object Lambda is used to transform data as it is retrieved from S3, not to reduce network egress costs or replace NAT gateway traffic. Option D is wrong because AWS Shield Advanced is a DDoS protection service that does not affect data transfer costs or routing between VPC and S3.

916
MCQhard

A financial services firm runs a stateless containerized trading dashboard on Amazon ECS with the Fargate launch type. The dashboard queries a backend over HTTPS and must present responses in under 200 ms. During market open, traffic triples within a few minutes and latency spikes because tasks take time to start. The team needs faster, more predictable scaling and wants to avoid over-provisioning during quiet periods. Which solution meets these requirements?

A.Switch the ECS service to the EC2 launch type with a large Auto Scaling group and enable cluster auto scaling.
B.Configure a target tracking scaling policy on the ECS service using the ALBRequestCountPerTarget metric with a longer cooldown period.
C.Configure a target tracking scaling policy on the ECS service using the ALBRequestCountPerTarget metric and set a short scale-out cooldown, keeping the minimum task count at a level that handles baseline traffic.
D.Use AWS Application Auto Scaling with a scheduled scaling action and a target tracking policy on CPU utilization, plus a warm pool of pre-initialized tasks.
AnswerC

Target tracking with ALBRequestCountPerTarget scales on the actual demand signal seen by the load balancer, so tasks are added as requests rise rather than after CPU saturates. A short scale-out cooldown lets the service respond within minutes, and a sensible minimum task count covers baseline traffic while avoiding over-provisioning during quiet periods.

Why this answer

Application Auto Scaling target tracking on a request-count-per-target metric reacts to incoming demand rather than lagging CPU, which suits a burst at market open. Keeping scale-out cooldowns short and setting a modest minimum task count balances responsiveness with cost. The other approaches either delay scaling, add slower instance-based capacity, or rely on signals that fire too late.

Exam trap

The trap here is treating a longer cooldown as a stabilizing improvement when it actually delays the scale-out needed to protect latency.

917
MCQeasy

A service role has an IAM policy granting kms:Decrypt for a specific AWS KMS key. The application still fails to decrypt with an AccessDenied error. What change most directly fixes this when the KMS key policy is missing the role’s permissions?

A.Update the KMS key policy to allow kms:Decrypt for the service role principal (or the assumed-role principal identity that the KMS key evaluates).
B.Add an IAM policy statement allowing s3:GetObject for the bucket that stores the encrypted data.
C.Enable a CloudFront distribution for the KMS key alias.
D.Create a VPC gateway endpoint for KMS to route decryption requests privately.
AnswerA

KMS authorization is controlled by the KMS key policy in addition to (not instead of) IAM identity policies. If the key policy does not allow the principal, KMS will deny kms:Decrypt even if the IAM policy allows it.

Why this answer

The AccessDenied error occurs because the KMS key policy does not grant the service role (or its assumed-role principal) permission to call kms:Decrypt. Even if the IAM policy attached to the role allows kms:Decrypt, KMS requires that the key policy explicitly authorize the principal (or the role’s assumed-role session) when the key policy is the sole authorization mechanism. Updating the key policy to include the service role principal (or the assumed-role ARN) with kms:Decrypt directly resolves the missing permission.

Exam trap

The trap here is that candidates assume IAM policies alone are sufficient for KMS authorization, but KMS key policies are resource-based and must explicitly include the principal (or the assumed-role session) when the key policy is the sole authorization mechanism.

How to eliminate wrong answers

Option B is wrong because s3:GetObject for the S3 bucket is unrelated to the KMS decryption failure; the error is specifically about KMS authorization, not S3 access. Option C is wrong because enabling a CloudFront distribution for the KMS key alias does not grant decryption permissions; CloudFront is a content delivery service and does not interact with KMS key policies. Option D is wrong because a VPC gateway endpoint for KMS only affects network routing for KMS API calls, not the IAM or key policy authorization; it does not grant the required kms:Decrypt permission.

918
MCQeasy

A company runs a stateless web application on Amazon EC2 instances behind an Application Load Balancer. Traffic has grown, and the operations team notices that individual instances are often underutilized while others are saturated because traffic is not evenly distributed. The team wants the load balancer to distribute requests more evenly across healthy targets. Which action should the team take?

A.Increase the idle timeout on the Application Load Balancer.
B.Change the load balancer from Application Load Balancer to Network Load Balancer.
C.Configure the target group to use the least outstanding requests routing algorithm.
D.Enable sticky sessions (session affinity) on the target group.
AnswerC

The least outstanding requests algorithm sends each new request to the target with the fewest in-flight requests, which naturally evens out load when request durations vary. It is well suited to stateless HTTP workloads and prevents slow or busy instances from accumulating a queue while idle targets sit unused, improving overall balance across healthy targets.

Why this answer

Even distribution of HTTP requests depends on the target group's routing algorithm. Least outstanding requests routes each new request to the target with the fewest active requests, which smooths out variance in request processing time. Round-robin alone can pile requests onto a slow target, so the least outstanding requests option is the right lever for a stateless workload.

Exam trap

The trap here is treating session affinity as a way to balance load, when it actually pins clients to single targets and worsens imbalance.

919
MCQmedium

A media company serves on-demand video to viewers worldwide from an Amazon S3 bucket in us-east-1. Viewers in Asia and Europe report slow start times because the first byte takes several seconds to arrive. The videos are already stored as objects and must remain in the us-east-1 bucket as the origin. Which solution improves global read performance with the LEAST operational effort?

A.Enable S3 Cross-Region Replication to replicate objects into buckets in ap-southeast-1 and eu-west-1, then point each region's viewers at its local bucket.
B.Move the bucket to an S3 Multi-Region Access Point and configure the application to read through the access point.
C.Create an Amazon CloudFront distribution with the S3 bucket as the origin and serve objects through the distribution.
D.Enable S3 Transfer Acceleration on the bucket and rewrite the application to upload and download through the accelerated endpoint.
AnswerC

CloudFront caches objects at edge locations near viewers and uses the AWS global backbone to fetch from the us-east-1 origin, so first-byte latency drops for Asia and Europe without duplicating data or changing the bucket. Because it is a managed service configured against the existing origin, operational effort is minimal and the objects stay where they are.

Why this answer

CloudFront is the managed content delivery layer for S3 origins: it caches objects at edge locations close to viewers and retrieves uncached content over the AWS backbone, cutting first-byte latency for globally distributed users. It needs no data duplication and no application rewrite, so it satisfies both the performance goal and the least-effort requirement while keeping the origin bucket in us-east-1.

Exam trap

The trap here is assuming that replicating or accelerating data movement is the same as caching content close to viewers, when only an edge cache removes the repeated long-haul fetch.

920
MCQmedium

A company uses AWS Organizations and has separate development, test, and production accounts. The security team wants to ensure that no one in the sandbox organizational unit can disable AWS CloudTrail or delete the central audit bucket, even if an account administrator creates permissive IAM policies later. Which control should they use?

A.Attach an identity-based policy in each account that denies CloudTrail changes.
B.Use a service control policy on the sandbox organizational unit to deny the prohibited actions.
C.Create an S3 bucket policy that allows only the audit team role to delete objects.
D.Apply a permission boundary to each IAM user in the sandbox accounts.
AnswerB

Service control policies are the correct governance mechanism for setting guardrails across multiple accounts in AWS Organizations. An SCP can explicitly deny sensitive actions such as disabling CloudTrail or deleting the audit bucket, and those denies apply even if administrators create local IAM policies that would otherwise allow the actions. SCPs do not grant permissions by themselves; they only constrain what account principals can ever do within the OU.

Why this answer

Service control policies (SCPs) are the correct mechanism because they act as a centralized guardrail at the AWS Organizations level, setting maximum permissions for all accounts in an organizational unit (OU). Even if an account administrator creates permissive IAM policies later, an SCP that explicitly denies disabling CloudTrail or deleting the central audit bucket will override those permissions, ensuring the security team's requirements are enforced across the sandbox OU.

Exam trap

The trap here is that candidates often confuse service control policies with IAM permission boundaries or resource-based policies, thinking that a bucket policy or permission boundary can prevent service-level actions like disabling CloudTrail, when only an SCP can enforce such restrictions across all principals in an entire OU.

Why the other options are wrong

A

Identity-based policies can be overridden by a more permissive policy attached by an account administrator, so they do not provide a guaranteed guardrail across all accounts in the organizational unit.

C

An S3 bucket policy only controls access to the bucket itself, not the ability to disable CloudTrail or delete the bucket from other services like IAM or CloudTrail. It does not prevent an account administrator from disabling CloudTrail entirely or deleting the bucket via the console or API.

D

Permission boundaries limit the maximum permissions for IAM users but do not prevent account administrators from creating permissive IAM policies that bypass the boundary, nor do they protect against actions taken by the root user or other principals. They are not effective for preventing CloudTrail disabling or bucket deletion across all users in an account.

When would these options actually be correct?

A

This option would be correct if the question asked for a way to restrict CloudTrail changes for a specific IAM user or role within a single account, without needing to enforce the restriction across multiple accounts or prevent override by an account admin.

C

This would be correct in a scenario where the requirement is to restrict deletion of objects within a specific S3 bucket to only an audit team role, while other users can still read or write. For example, a question asking 'How to ensure only the audit team can delete audit logs from the central bucket?'

D

A question where the requirement is to restrict the maximum permissions for specific IAM users or roles within an account, such as limiting developers to read-only access while allowing them to create their own IAM roles within those bounds.

Why candidates pick the wrong answer

A

Candidates may think identity-based policies are sufficient for restricting actions, but they overlook that account administrators can attach permissive policies that override the deny, making SCPs necessary for organization-wide enforcement.

C

Candidates may think a bucket policy is sufficient to protect the audit bucket, overlooking that CloudTrail can be disabled independently or that the bucket itself can be deleted via other means.

D

Candidates may confuse permission boundaries with service control policies, thinking they can centrally enforce restrictions on all users, but boundaries only apply to IAM principals and can be overridden by account administrators.

921
MCQmedium

Your mobile app writes events to a single DynamoDB table with partition key = customerId and sort key = eventTime. During a promotional campaign, one tenant ("ACME") generates far more traffic than others. CloudWatch shows sustained throttling (ProvisionedThroughputExceeded) and elevated p99 latency only for that tenant. The workload pattern cannot be changed to a completely different schema, but you can change how items are partitioned. Which design change is most likely to reduce the hot-partition throttling while keeping efficient reads for ACME?

A.Use the same partition key (customerId), but increase the table’s provisioned capacity for that tenant.
B.Change the partition key to a salted key such as customerId + shard number, and include the eventTime ordering using the sort key.
C.Switch to on-demand capacity mode and keep the partition key unchanged.
D.Enable Global Tables so that reads are served from a nearby replica for ACME.
AnswerB

Hot-partition throttling happens when a single logical partition (one partition key value) receives more requests than it can serve. By salting the partition key (for example, customerId#shardId), ACME’s writes are spread across multiple physical partitions, reducing request rate per partition and lowering throttling. Efficient reads for ACME can be preserved by querying only the shard partitions that belong to ACME (for example, using a small, deterministic set of shardIds and issuing parallel queries per shard, then merging results). This avoids scanning the whole table and keeps access patterns predictable while improving tail latency.

Why this answer

Salting the partition key by appending a shard number (e.g., customerId + random digit) distributes ACME's writes across multiple partitions, eliminating the hot partition. The sort key still preserves eventTime ordering, so queries for a specific customer can be parallelized across shards and merged client-side or via a composite sort key pattern, maintaining efficient reads.

Exam trap

The trap here is that candidates assume increasing capacity or switching to on-demand alone solves hot partitions, but they overlook DynamoDB's fixed per-partition throughput limits that require key design changes to distribute load.

How to eliminate wrong answers

Option A is wrong because increasing provisioned capacity for a single tenant does not solve the hot-partition issue; DynamoDB distributes capacity across partitions, and a single partition's throughput is capped at 3,000 RCU or 1,000 WCU regardless of table-level settings. Option C is wrong because switching to on-demand capacity mode only handles traffic spikes at the table level, but a single hot partition still hits the same per-partition throughput limits (3,000 RCU/1,000 WCU), causing throttling. Option D is wrong because Global Tables replicate data across regions for low-latency reads and disaster recovery, but they do not redistribute write load within a single table; ACME's writes still target the same partition key in the source region, so throttling persists.

922
MCQmedium

An Aurora PostgreSQL cluster is experiencing high read latency because 85% of traffic consists of read-only queries. The write workload must stay on the writer instance, and the team wants to offload reads without changing the application’s core query patterns. What is the best architectural option?

A.Increase the writer instance size so it can handle more reads and writes simultaneously.
B.Add Aurora reader instances (read replicas) and route read queries to the reader endpoint while keeping writes on the writer endpoint.
C.Enable Multi-AZ failover only and rely on the standby to serve reads in normal operation.
D.Move the read workload to ElastiCache Redis while keeping DynamoDB as the SQL data source.
AnswerB

Aurora reader instances are designed for exactly this pattern: they provide dedicated compute capacity for read-only workloads. By sending read queries to the reader endpoint and keeping writes on the writer endpoint, the cluster can scale read performance without forcing reads to contend with write processing on the writer.

Why this answer

Adding Aurora reader instances (read replicas) and routing read queries to the reader endpoint offloads read traffic from the writer instance without altering application query patterns. Aurora reader endpoints automatically distribute read-only connections across all replicas, reducing latency on the writer while keeping writes on the writer instance. This directly addresses the 85% read-heavy workload without requiring application changes.

Exam trap

The trap here is that candidates often confuse Multi-AZ standby instances (which are passive and cannot serve reads) with Aurora reader replicas (which are active and can serve reads), leading them to incorrectly select Option C.

Why the other options are wrong

A

Increasing the writer instance size does not offload reads; it only scales the single writer instance, which still handles all read traffic and does not reduce read latency from high read concurrency.

C

In Aurora, a Multi-AZ standby (writer failover target) does not serve read traffic; it is only used for failover. To offload reads, you need dedicated reader instances with a separate reader endpoint.

D

Option D is wrong because it suggests using ElastiCache Redis with DynamoDB as the SQL data source, but the question specifies an Aurora PostgreSQL cluster. DynamoDB is a NoSQL database, not a SQL data source, and this approach would require significant application changes to query patterns, violating the constraint of not changing core query patterns.

When would these options actually be correct?

A

This option would be correct if the question stated that the cluster is CPU-bound on the writer due to a mix of reads and writes, and the application cannot tolerate any read replica lag or connection routing changes, requiring a vertical scaling approach.

C

For a non-Aurora RDS database (e.g., RDS for MySQL or PostgreSQL) where you need high availability and want to offload read traffic to a standby, enabling Multi-AZ with the 'standby can serve reads' option (if supported) would be correct. The question would specify a single-instance RDS DB with Multi-AZ and the need to use the standby for reads.

D

This option would be correct in a scenario where the application uses a NoSQL database like DynamoDB as its primary data store, experiences high read traffic, and can tolerate eventual consistency. The team wants to offload reads without altering query patterns, and using ElastiCache Redis as a caching layer in front of DynamoDB would reduce read latency.

Why candidates pick the wrong answer

A

Candidates may think that a larger instance can handle more total throughput, overlooking that read replicas are specifically designed to offload read traffic and reduce latency for read-heavy workloads.

C

Candidates may confuse Aurora's Multi-AZ with RDS Multi-AZ, or incorrectly assume that a standby can handle read requests in normal operation, similar to read replicas.

D

Candidates may choose this option because they recognize that caching (ElastiCache) is a common solution for read-heavy workloads, and they might overlook the specific database type (Aurora PostgreSQL vs. DynamoDB) or assume that any caching layer can be seamlessly integrated without considering the underlying data store compatibility.

923
MCQeasy

A media startup stores user-uploaded video files in an Amazon S3 bucket in the us-east-1 Region. The compliance team requires that the data remain recoverable if an entire AWS Region becomes unavailable, and that recovery can be performed by pointing applications at a different endpoint. Cost should be minimized while still meeting the requirement. Which solution should a solutions architect recommend?

A.Configure an S3 Lifecycle rule to transition objects to S3 Glacier Deep Archive after 30 days for regional durability.
B.Enable S3 Cross-Region Replication to a bucket in a second Region and configure the application to use the replica bucket endpoint during a regional failure.
C.Enable S3 Transfer Acceleration on the bucket so uploads and downloads are faster during a regional event.
D.Enable S3 Versioning on the existing bucket and rely on version history to restore objects if the Region fails.
AnswerB

Cross-Region Replication copies objects asynchronously to a bucket in another Region, giving a durable copy that survives a full regional outage. Because replication is object-level and pay-as-you-go, cost stays lower than continuously active multi-Region compute, and the application can be repointed to the replica bucket's endpoint during failover.

Why this answer

Protecting S3 data against a full regional outage requires a copy in a different Region, which Cross-Region Replication provides asynchronously and cost-effectively. Versioning, lifecycle transitions, and Transfer Acceleration all operate within the source Region and therefore cannot deliver recovery when that Region is lost.

Exam trap

The trap here is confusing durability features that live inside one Region, such as versioning or Glacier transitions, with actual cross-Region resilience.

924
MCQhard

Based on the exhibit, the current disaster recovery design misses the RTO target even though the database replica is current. Which deployment model best meets the requirements with the least always-on cost?

A.Pilot light, because only the database needs to be running in the secondary Region.
B.Warm standby, because a scaled-down application stack stays running in the secondary Region and can take over faster.
C.Active-active, because both Regions should always serve traffic to guarantee the RTO.
D.Backup and restore, because restoring from backups is the least expensive DR model available.
AnswerB

Warm standby is the best fit when you need faster recovery than pilot light but do not want the cost of full active-active capacity. The exhibit shows that starting the application stack from zero consumes most of the recovery time. Keeping a reduced but functional stack running in the secondary Region removes that startup delay and should bring the total recovery time within the 15-minute RTO while still keeping always-on cost below full production duplication.

Why this answer

Warm standby is the correct choice because it keeps a scaled-down application stack running in the secondary Region, which can be scaled up quickly to handle production traffic. This design meets the RTO target by reducing failover time compared to a pilot light, while avoiding the higher always-on cost of an active-active deployment.

Exam trap

The trap here is that candidates confuse pilot light with warm standby, assuming that only the database needs to be running to meet the RTO, but they overlook the time required to provision the application stack on failover.

Why the other options are wrong

A

The pilot light model only keeps the database running in the secondary Region, not the application stack. This means the application must be provisioned and scaled up after a disaster, which takes too long to meet the RTO, even if the database is current.

C

Active-active requires both regions to serve live traffic continuously, which incurs higher always-on costs than warm standby. The question asks for the 'least always-on cost' while meeting RTO, and active-active is more expensive because it runs full production capacity in both regions.

D

Backup and restore typically has a high RTO because it involves restoring data from backups, which is slower than having a running replica. The question states the database replica is current, so a warm standby with a scaled-down application stack can meet the RTO faster.

When would these options actually be correct?

A

A question where the RTO is longer (e.g., hours) and the primary concern is minimizing cost while keeping critical data available. For example: 'A company needs a DR strategy that minimizes cost but can restore database access within 4 hours. The application can be rebuilt quickly from scripts.'

C

An exam scenario where the application requires near-zero RTO (e.g., under 1 minute) and can tolerate the higher cost, such as a global e-commerce platform that must remain available during a regional outage without any traffic rerouting delay.

D

A question where the RTO is lenient (e.g., hours) and the primary goal is to minimize cost, with no requirement for a current database replica. For example: 'Which DR model is the least expensive and can tolerate an RTO of several hours?'

Why candidates pick the wrong answer

A

Candidates may think that because the database replica is current, only the database needs to be running to meet the RTO, overlooking the time required to start and configure the application stack.

C

Candidates may think active-active is the fastest failover model and assume it always meets RTO best, overlooking the cost constraint and that warm standby can achieve the same RTO at lower cost.

D

Candidates may think backup and restore is the cheapest option and assume it can meet any RTO if backups are frequent, overlooking the time needed to restore and the fact that a current replica already exists.

925
MCQmedium

A team wants to remove a bastion host used for administrative access to EC2 instances in private subnets. The instances should be reachable only for occasional troubleshooting by engineers who authenticate with AWS SSO. What is the best secure alternative within AWS, assuming the instances already have an instance profile attached?

A.Use AWS Systems Manager Session Manager, enabling the required SSM permissions in the instance profile and restricting access to engineers via IAM.
B.Keep the bastion host but move it into a private subnet; engineers can connect by using a corporate VPN into the VPC.
C.Attach a public IP to each private instance so engineers can SSH directly and use security groups to restrict access.
D.Create a security group rule that allows engineers’ source IP addresses to reach instances over RDP on port 3389.
AnswerA

Session Manager avoids inbound SSH from the internet by initiating interactive sessions through Systems Manager. The instance profile must allow SSM actions like StartSession, and engineers’ IAM permissions restrict who can connect. This is a commonly recommended bastion-free alternative that improves security and reduces exposed network paths.

Why this answer

AWS Systems Manager Session Manager provides secure, auditable, agent-based access to EC2 instances without requiring a bastion host, public IPs, or open inbound ports. Since the instances already have an instance profile, you only need to add the required SSM permissions (e.g., AmazonSSMManagedInstanceCore) to that profile and use IAM policies to restrict Session Manager access to engineers authenticated via AWS SSO. This eliminates the bastion host while maintaining secure, on-demand troubleshooting access.

Exam trap

The trap here is that candidates often think a bastion host is the only way to access private instances, overlooking that AWS Systems Manager Session Manager provides a fully managed, agent-based alternative that eliminates the need for any bastion host or open inbound ports.

How to eliminate wrong answers

Option B is wrong because moving the bastion host to a private subnet and using a corporate VPN still leaves a persistent bastion host that must be patched and managed, and it does not eliminate the attack surface or the need for SSH/RDP key management. Option C is wrong because attaching a public IP to each private instance directly exposes them to the internet, violating the principle of least privilege and increasing the attack surface, even with security group restrictions. Option D is wrong because allowing engineers’ source IPs over RDP port 3389 requires opening inbound ports and managing IP whitelists, which is less secure than agentless access and does not integrate with AWS SSO for authentication.

926
MCQeasy

A retail platform needs disaster recovery across AWS Regions. The business requirement is: RTO up to 6 hours, RPO up to 1 hour, and they want the ability to start serving quickly during a Region outage but do not want to run full production capacity continuously. Which DR strategy best fits these requirements?

A.Backup and restore only, with no continuously running infrastructure in the secondary Region.
B.Pilot light, keeping only the minimum resources needed to bootstrap the environment.
C.Warm standby, keeping a reduced but ready-to-scale environment in the secondary Region.
D.Multi-site active-active, serving production traffic from both Regions at all times.
AnswerC

Warm standby in the secondary Region runs a scaled-down but fully functional copy of your production stack, typically with key databases and services already deployed and synchronized. You can provision extra compute capacity on demand via Auto Scaling or pre-provisioned cluster resizing to reach full production load within the 6-hour RTO. This balances cost and recovery speed by keeping idle but ready infrastructure that can be quickly scaled up, making it the most appropriate choice for the stated requirement.

Why this answer

Warm standby is the correct strategy because it maintains a scaled-down but fully functional copy of the production environment in the secondary Region, which can be scaled up within the 6-hour RTO. The RPO of 1 hour is met by continuous replication (e.g., Amazon RDS cross-Region read replicas or DynamoDB global tables), and the reduced footprint avoids the cost of full production capacity while still enabling rapid failover.

Exam trap

The trap here is that candidates confuse pilot light with warm standby, assuming that any minimal running infrastructure qualifies as pilot light, but warm standby specifically requires a scaled-down but fully functional environment that can serve traffic immediately after scaling, whereas pilot light requires significant provisioning before it can serve traffic.

Why the other options are wrong

A

Backup and restore typically has RPOs of hours or days and RTOs of 24+ hours, failing to meet the 1-hour RPO and 6-hour RTO requirements.

B

The pilot light strategy typically has RTO of 10-15 minutes and RPO of a few minutes, which is faster than the required 6-hour RTO and 1-hour RPO, but it does not meet the requirement to 'start serving quickly' during a Region outage because it requires provisioning and scaling resources after failover, leading to longer recovery time than warm standby.

D

Multi-site active-active requires running full production capacity in both Regions continuously, which contradicts the requirement to not run full production capacity continuously.

When would these options actually be correct?

A

For a non-critical application with RTO > 24 hours and RPO > 24 hours, where cost is the primary concern and data loss of up to a day is acceptable.

B

A scenario where the business requires very low cost for the DR site, can tolerate longer RTO (e.g., 12-24 hours), and has minimal RPO requirements (e.g., 1 hour). For example, a non-critical internal tool that can afford significant downtime during a disaster.

D

An exam question requiring zero RTO and zero RPO for a mission-critical application with unlimited budget, where continuous active-active traffic distribution is acceptable.

Why candidates pick the wrong answer

A

Candidates may think backup and restore is the simplest and cheapest DR method, overlooking the strict RPO/RTO requirements in the question.

B

Candidates may confuse 'pilot light' with 'warm standby' because both involve running some resources in the secondary Region, but they underestimate the additional provisioning time needed for pilot light to become fully operational, making it unsuitable for the 'quickly' requirement.

D

Candidates may think active-active provides the fastest failover and meets RTO/RPO, but overlook the cost and capacity requirement that conflicts with the 'not run full production capacity continuously' constraint.

927
MCQmedium

An event ingestion service writes to a DynamoDB table where the partition key is tenantId and the sort key is eventTime. During a campaign, one tenant generates a disproportionate share of traffic, causing write throttling and increased latency for that tenant’s writes. You can change the data model and application queries, but you must still efficiently retrieve events for a tenant for the last 10 minutes. Which change best improves write throughput by reducing hot partitions?

A.Keep tenantId as the partition key and rely on DynamoDB adaptive capacity to automatically remove all throttling.
B.Add a shard attribute to the partition key (partition key = tenantId#shard, where shard is randomly selected from a fixed range). Query all shards for the tenant for eventTime values in the last 10 minutes, then merge results in the application.
C.Change the sort key to eventTimeBucket (for example, eventTime rounded to 1-minute buckets) while keeping the partition key as tenantId.
D.Enable DAX and use it for write operations so throttled writes are served from cache instead of reaching DynamoDB.
AnswerB

This “write sharding” spreads a tenant’s traffic across multiple partition key values, which distributes the write load across multiple DynamoDB partitions (and thus multiple throughput slices). Reads for the last 10 minutes remain efficient because each shard still supports a sort-key range query on eventTime; the application merges results across shards.

Why this answer

It distributes writes for a hot tenant across multiple partitions by appending a random shard suffix to the tenantId partition key. This eliminates a single hot partition, allowing DynamoDB to scale write capacity horizontally. The application can then query all shards for the last 10 minutes and merge results, satisfying the retrieval requirement.

Exam trap

The trap here is that candidates assume adaptive capacity or caching (DAX) can solve write throttling, but neither addresses the root cause—a single partition exceeding its write capacity—which requires redistributing the partition key across multiple physical partitions.

How to eliminate wrong answers

Option A is wrong because DynamoDB adaptive capacity can only mitigate moderate hot spots by temporarily allocating extra capacity, but it cannot eliminate throttling when a single partition exceeds its 1,000 WCU or 3,000 RCU limit; sustained high traffic from one tenant will still cause throttling. Option C is wrong because changing the sort key to eventTimeBucket does nothing to distribute writes across partitions—the partition key remains tenantId, so all writes for that tenant still target the same partition, leaving the hot partition problem unsolved. Option D is wrong because DAX is a read-through cache and does not handle write operations; throttled writes are not served from cache, and DAX cannot increase write throughput or reduce hot partition contention.

928
MCQeasy

A containerized service needs to read exactly one secret value from AWS Secrets Manager. The secret’s ARN is already known, and the secret is encrypted with the AWS-managed KMS key for Secrets Manager, so no separate KMS permissions are needed for this question. The service does not need to list secrets, create secrets, rotate them, or write updates. What is the most least-privilege IAM permission statement to grant the service role?

A.Allow secretsmanager:GetSecretValue on the specific secret ARN only.
B.Allow secretsmanager:* on all resources in the account.
C.Allow secretsmanager:ListSecrets so the service can discover the secret ARN at runtime.
D.Allow secretsmanager:PutSecretValue so the service can retrieve and update the secret value.
AnswerA

For a read-only use case where the secret ARN is already known, the minimum required Secrets Manager action is secretsmanager:GetSecretValue. Scoping the resource to only that secret ARN minimizes blast radius if the role is compromised.

Why this answer

The service only needs to read a single secret value, and the least-privilege permission is to allow only the `secretsmanager:GetSecretValue` action on that specific secret's ARN. This grants exactly the required read access without any additional capabilities, adhering to the principle of least privilege. Since the secret is encrypted with the AWS-managed KMS key for Secrets Manager, no separate KMS permissions are needed, as the key policy automatically grants access to the Secrets Manager service.

Exam trap

The trap here is that candidates often choose a broader permission like `secretsmanager:*` or `secretsmanager:ListSecrets` because they confuse the need to discover the secret with the need to read it, or they overlook that the ARN is already known, making list actions unnecessary.

How to eliminate wrong answers

Option B is wrong because `secretsmanager:*` on all resources grants full administrative access to all secrets in the account, which is far more permissive than needed and violates least privilege. Option C is wrong because `secretsmanager:ListSecrets` allows listing all secret names and ARNs in the account, which is unnecessary since the secret ARN is already known, and it provides no ability to read the secret value itself. Option D is wrong because `secretsmanager:PutSecretValue` allows updating the secret value, which is not required and introduces unnecessary write permissions that could lead to accidental or malicious modification.

929
MCQeasy

A media company stores 50 TB of finalized video masters in Amazon S3 that must be retained for seven years for regulatory compliance. The files are accessed only during occasional legal audits, roughly once every two years, and retrieval latency of several hours is acceptable. The company wants the LOWEST possible storage cost while preserving durability. Which storage class should they choose?

A.S3 Glacier Flexible Retrieval
B.S3 Standard-Infrequent Access
C.S3 One Zone-IA
D.S3 Glacier Deep Archive
AnswerD

S3 Glacier Deep Archive is designed for long-term retention of data accessed less than once per year, offering the lowest storage cost of any S3 class while maintaining eleven nines of durability. Retrieval takes up to 12 hours, which fits the acceptable latency stated for rare legal audits. For 50 TB held seven years with almost no access, this class minimizes spend without sacrificing durability.

Why this answer

Data accessed less than once per year and tolerant of hours-long retrieval belongs in S3 Glacier Deep Archive, which offers the lowest per-GB storage price in S3 while retaining eleven nines of durability. Glacier Flexible Retrieval and the IA classes cost more per GB and target more frequent access patterns, so they would increase seven-year spend without adding value.

Exam trap

The trap here is reaching for a mid-tier archive class or an Infrequent Access class out of habit, when the stated access frequency of once every two years and tolerance for hours of latency points squarely at the cheapest archive tier.

930
MCQeasy

A security team needs an audit trail to investigate suspicious API activity across multiple AWS accounts. Which AWS approach best provides centralized visibility into who did what, when, for service API calls?

A.Create an AWS CloudTrail organization trail that delivers logs to a centralized, access-controlled S3 bucket.
B.Enable AWS Config only for EC2 security groups and rely on it for API call auditing.
C.Turn on S3 server access logging for every bucket and assume it covers all AWS services.
D.Use only Amazon CloudWatch alarms with no logging destination to reduce storage costs.
AnswerA

An AWS Organizations organization trail centralizes management and API activity logs across accounts. CloudTrail provides detailed event records including the requesting principal, source information, event time, and the specific API action, which supports forensic investigation.

Why this answer

AWS CloudTrail organization trail is the correct approach because it captures all management and data events across multiple AWS accounts within an AWS Organizations structure, delivering them to a single, centralized S3 bucket. This provides a unified, immutable audit trail of who performed which API call, when, and from which source IP, enabling the security team to investigate suspicious activity with full visibility. The centralized bucket can be access-controlled with S3 bucket policies and IAM to ensure only authorized personnel can view the logs.

Exam trap

The trap here is that candidates often confuse AWS Config (which tracks resource configuration changes) with CloudTrail (which records API calls), leading them to pick Option B, but Config does not provide the who, what, when details needed for an API audit trail.

How to eliminate wrong answers

Option B is wrong because AWS Config is a resource inventory and configuration change tracking service, not an API call auditor; it does not capture who made the API call or the full request/response details. Option C is wrong because S3 server access logging only records requests made to S3 buckets, not API calls for other AWS services like EC2, IAM, or Lambda, leaving a massive gap in the audit trail. Option D is wrong because CloudWatch alarms only trigger on metric thresholds and do not store or provide any log data for forensic investigation; without a logging destination, there is no audit trail at all.

931
MCQmedium

Based on the exhibit, what should the security team implement so developers can create AWS Lambda execution roles, but no developer-created role can ever exceed the approved permission set?

A.Place the developers in an IAM group with a deny-only managed policy attached.
B.Require a permissions boundary on every developer-created role and set the boundary to the approved maximum permissions.
C.Use an AWS Organizations SCP to grant only the approved Lambda permissions directly to the developer roles.
D.Create the roles with inline policies only, because inline policies are always safer than managed policies.
AnswerB

A permissions boundary limits the highest permissions a role can ever have, even if someone attaches broader policies later. This is the right guardrail when developers are allowed to create roles but must stay within a security-approved ceiling. It still lets them work independently while preventing privilege escalation through policy attachment.

Why this answer

A permissions boundary explicitly defines the maximum permissions that an IAM role can have, and when attached to developer-created roles, it prevents any role from exceeding the approved set of permissions, even if the developer attaches a more permissive policy. This directly addresses the requirement that no developer-created role can ever exceed the approved permission set, as the boundary acts as a hard cap.

Exam trap

The trap here is that candidates often confuse permissions boundaries with SCPs or assume that inline policies are more restrictive, but the key is that a permissions boundary is the only mechanism that directly caps the maximum permissions of a specific role without affecting other principals.

Why the other options are wrong

A

A deny-only managed policy would block all actions, not just limit permissions to an approved set. It does not allow developers to create roles with any permissions, and it cannot be used to set a maximum permission boundary.

C

An SCP cannot grant permissions; it only denies or allows permissions at the account level. It cannot be used to grant specific Lambda permissions directly to developer roles, and it does not enforce a maximum permission boundary on individual roles.

D

Inline policies are not inherently safer than managed policies; they are attached directly to a role and can still exceed approved permissions. The question requires a mechanism to enforce a maximum permission set, which inline policies cannot guarantee.

When would these options actually be correct?

A

This option would be correct if the question asked for a method to explicitly deny specific actions (e.g., deny access to certain AWS services) for all developers, without needing to manage individual policies.

C

In a scenario where the organization needs to centrally restrict all IAM actions across multiple accounts, such as preventing any role from accessing certain high-risk services, an SCP would be the correct tool to apply a broad deny policy at the organizational unit level.

D

When the question asks for the most secure way to attach permissions to a role that should never be shared or reused across multiple roles, and the concern is about limiting the scope of permissions to a single role without risk of unintended attachment to other roles.

Why candidates pick the wrong answer

A

Candidates may think a deny-only policy is a simple way to restrict permissions, misunderstanding that it would prevent all actions rather than setting a maximum allowed set.

C

Candidates may confuse SCPs with permission boundaries, thinking SCPs can also set granular maximum permissions on roles, when in fact SCPs are account-wide guardrails and cannot be attached to individual roles.

D

Candidates may believe inline policies are safer because they are tightly coupled to a role and less likely to be accidentally attached elsewhere, but they overlook that inline policies can still grant excessive permissions and are harder to audit.

932
MCQmedium

Your team hosts a private web app on an S3 bucket and serves it through CloudFront using a modern Origin Access Control (OAC). After deployment, users receive HTTP 403 from CloudFront with the S3 origin error "AccessDenied". Which S3 bucket policy change best aligns with CloudFront OAC so the distribution can fetch objects privately?

A.Allow the CloudFront service principal cloudfront.amazonaws.com to perform s3:GetObject, and scope access with a condition on AWS:SourceArn matching your CloudFront distribution ARN.
B.Allow only the S3 bucket owner account to perform s3:GetObject without any condition, so CloudFront can inherit access automatically.
C.Add a policy statement that denies s3:GetObject when the request does not include the header CloudFront-Viewer-Country.
D.Grant s3:GetObject permission to an Origin Access Identity (OAI) canonical user ID even though you are using Origin Access Control (OAC).
AnswerA

With CloudFront OAC, the request to S3 is authorized using the CloudFront service principal. Granting s3:GetObject to cloudfront.amazonaws.com and constraining it with AWS:SourceArn to the specific distribution is the standard secure pattern for private S3 origins.

Why this answer

CloudFront Origin Access Control (OAC) requires an explicit S3 bucket policy that allows the CloudFront service principal (`cloudfront.amazonaws.com`) to perform `s3:GetObject`, and the recommended best practice is to scope the permission using a condition on `AWS:SourceArn` matching the specific CloudFront distribution ARN. This ensures that only requests originating from that distribution can access the bucket objects, preventing unauthorized access from other sources.

Exam trap

The trap here is that candidates often confuse Origin Access Control (OAC) with the older Origin Access Identity (OAI) and incorrectly select an OAI-based policy (Option D), or they assume that bucket owner permissions automatically extend to CloudFront (Option B), failing to recognize that OAC requires an explicit service principal-based policy with a source ARN condition.

Why the other options are wrong

B

CloudFront OAC does not automatically inherit permissions from the bucket owner; it requires an explicit bucket policy that allows the CloudFront service principal with a condition on the source ARN. Option B lacks this condition and principal, so it would not grant CloudFront access.

C

This option is wrong because CloudFront OAC does not use headers like CloudFront-Viewer-Country for authentication; the 403 error is due to missing permissions for CloudFront to access S3, not due to missing headers.

D

The question specifies using Origin Access Control (OAC), not Origin Access Identity (OAI). OAI uses a canonical user ID, but OAC uses the CloudFront service principal with a source ARN condition. Granting permissions to an OAI canonical user ID does not work with OAC, so the distribution will still receive 403 errors.

When would these options actually be correct?

B

This option would be correct if the question asked about a scenario where the S3 bucket is publicly accessible and no CloudFront origin access control is needed, such as when serving static assets directly from S3 without CloudFront.

C

This option would be correct in a scenario where you need to restrict access to your S3 content based on viewer's country, such as when serving content only to specific countries for licensing or compliance reasons, and you want to deny access from other countries.

D

If the question stated that the distribution uses an Origin Access Identity (OAI) instead of OAC, then granting s3:GetObject permission to the OAI's canonical user ID would be correct. For example: 'You configured CloudFront with an OAI to restrict access to an S3 bucket. Which bucket policy allows CloudFront to fetch objects?'

Why candidates pick the wrong answer

B

Candidates may mistakenly think that CloudFront inherits the bucket owner's permissions automatically, not realizing that OAC requires an explicit policy statement with the CloudFront service principal and source ARN condition.

C

Candidates might think that adding a header-based condition could resolve the access issue, misunderstanding that the 403 is about authorization, not about missing headers for geo-restriction.

D

Candidates may confuse OAI with OAC, or think that any identity-based access works similarly. They might recall that OAI uses canonical user IDs and assume it applies to OAC as well, not realizing OAC uses a different authorization mechanism.

933
MCQeasy

A mobile app reads the same product details many times per minute from Amazon DynamoDB. The table design is already correct, but repeated reads are still causing noticeable latency. Which service should the team add to improve read performance?

A.Amazon DAX
B.Amazon EFS
C.AWS Lambda
D.Amazon SNS
AnswerA

Amazon DAX is a DynamoDB-compatible in-memory caching service that delivers up to 10x read performance improvement, reducing read latency from single-digit milliseconds to microseconds. It sits in front of your DynamoDB tables and transparently intercepts repeated read calls, serving them from the cache without additional application code. This makes DAX the correct choice for a mobile app that reads identical product details many times, as it directly targets the repeated-read bottleneck by bypassing the database's disk access and reducing the load on DynamoDB.

Why this answer

Amazon DAX (DynamoDB Accelerator) is an in-memory cache specifically designed for DynamoDB. It reduces read latency from single-digit milliseconds to microseconds by caching frequently accessed items, which directly addresses the repeated read pattern described in the question.

Exam trap

The trap here is that candidates may confuse DAX with ElastiCache (which is generic and not DynamoDB-native) or assume that Lambda or SNS can somehow accelerate reads, but DAX is the only service purpose-built for DynamoDB read acceleration.

Why the other options are wrong

B

Amazon EFS is a file storage service for EC2 instances, not a caching layer for DynamoDB. It cannot reduce read latency for DynamoDB queries because it does not integrate with DynamoDB's API or provide in-memory acceleration.

C

AWS Lambda is a compute service for running code in response to events, not a caching layer. It cannot reduce read latency for repeated DynamoDB reads because it does not cache data; adding Lambda would introduce additional invocation overhead, worsening latency.

D

Amazon SNS is a pub/sub messaging service for notifications and event-driven workflows, not a caching layer. It cannot reduce read latency for repeated DynamoDB reads because it does not store or serve data from a cache.

When would these options actually be correct?

B

A question where an application running on multiple EC2 instances needs a shared, scalable file system for storing and accessing common files (e.g., configuration files, logs, or media) across instances. EFS would be the correct choice for a shared NFS file system.

C

A question where a DynamoDB stream triggers a function to transform or aggregate data before storing it elsewhere (e.g., update an Amazon ElastiCache cluster or write to Amazon S3) would make Lambda the correct answer. For example: 'A team needs to process item changes from a DynamoDB table and update a materialized view in near real time.'

D

A correct scenario: 'An application needs to send real-time alerts to multiple subscribers (e.g., email, SMS, Lambda) whenever a new item is inserted into a DynamoDB table.' In that case, Amazon SNS would be the right service to fan out notifications.

Why candidates pick the wrong answer

B

Candidates might confuse EFS as a general-purpose 'cache' because it offers low-latency file access, but they overlook that it is not designed for database query caching and cannot be used with DynamoDB directly.

C

Candidates may think Lambda can 'speed up' reads by preprocessing or caching data, but Lambda is stateless and not designed for low-latency data retrieval. The temptation comes from overgeneralizing Lambda's role in optimizing data pipelines.

D

Candidates might think SNS can cache or buffer repeated reads because it's a managed service that can decouple components, but they confuse its notification role with caching or data serving capabilities.

934
MCQmedium

A media company stores original video assets in an Amazon S3 bucket in the us-east-1 Region. Editors in Europe report slow downloads, and the legal team requires that a copy of every asset exist in eu-west-1 within 15 minutes of upload, with the ability to fail over reads to the European copy during a Regional impairment. Which S3 feature should the architects enable?

A.S3 Multi-Region Access Points configured against the existing us-east-1 bucket, with no replication rule created.
B.S3 Object Lambda access points in eu-west-1 that transform objects on retrieval from the us-east-1 bucket.
C.S3 Cross-Region Replication with a replication rule that targets the eu-west-1 bucket and uses S3 Replication Time Control.
D.S3 Same-Region Replication to a second bucket in us-east-1, with an S3 Transfer Acceleration endpoint used by the European editors.
AnswerC

Cross-Region Replication copies new objects to the destination bucket automatically, and S3 Replication Time Control provides a service level agreement that most objects replicate within 15 minutes. This meets both the latency goal for European readers and the requirement for a failover copy in eu-west-1. It is the native, managed mechanism for the stated RPO.

Why this answer

The requirement is a durable, timely copy of each object in a second Region, which is exactly what Cross-Region Replication provides. Adding S3 Replication Time Control turns the timing expectation into a defined service level, aligning the solution with the 15-minute RPO the legal team specified. Other S3 features change routing or transform data but do not create the required replica.

Exam trap

The trap here is confusing features that change how clients reach S3 with features that actually copy object data to another Region.

935
MCQmedium

A Lambda function for a IoT ingestion API needs to read a database password. The password must rotate automatically every 30 days and should not be stored in environment variables. Which service should be used?

A.AWS Secrets Manager with rotation enabled
B.A KMS-encrypted Lambda environment variable
C.AWS Systems Manager Parameter Store SecureString without automation
D.An encrypted object in Amazon S3
AnswerA

AWS Secrets Manager is purpose-built for storing and retrieving secrets at runtime, and the rotation-enabled configuration automatically rotates database credentials on a schedule using an accompanying AWS Lambda function. This ensures the Lambda function always reads a valid, current secret without manual intervention, and access can be scoped via IAM policies. It also integrates with CloudTrail for auditing secret access, making it the appropriate choice for a production IoT ingestion pipeline.

Why this answer

AWS Secrets Manager is the correct choice because it is purpose-built for securely storing, automatically rotating, and managing secrets such as database passwords. With rotation enabled, Secrets Manager can automatically rotate the password every 30 days without requiring custom code, and it integrates natively with Lambda via the AWS SDK to retrieve the secret at runtime, avoiding storage in environment variables.

Exam trap

The trap here is that candidates often confuse AWS Systems Manager Parameter Store SecureString with Secrets Manager, but Parameter Store lacks native automatic rotation, making it unsuitable for the 30-day rotation requirement without additional custom automation.

How to eliminate wrong answers

Option B is wrong because storing a KMS-encrypted password in a Lambda environment variable still exposes the encrypted value in the function's configuration and does not support automatic rotation; the password would need manual rotation and re-deployment. Option C is wrong because AWS Systems Manager Parameter Store SecureString without automation does not provide automatic rotation; it only stores the secret securely, requiring custom logic to rotate the password every 30 days. Option D is wrong because an encrypted object in Amazon S3 lacks native rotation capabilities and introduces unnecessary complexity for secret retrieval, as Lambda would need to decrypt the object and manage rotation manually.

Page 12

Page 13 of 13