Courseiva

SAA-C03 (SAA-C03) — Questions 826–900

935 questions total · 13pages · All types, answers revealed

Page 11

Page 12 of 13

Page 13
826
MCQhard

A trading platform ingests market data events at very high volume and must deliver them with the lowest possible latency to multiple independent consumer applications. Each consumer must read the full stream independently, and ordering must be preserved per instrument symbol. Which solution meets these requirements?

A.Create an Amazon SQS FIFO queue with a message group ID of the instrument symbol, and have each consumer application use a separate queue.
B.Use AWS Database Migration Service with a change data capture task to replicate events into an Amazon Aurora cluster that each consumer queries.
C.Create an Amazon SNS topic and subscribe each consumer application with an Amazon SQS standard queue endpoint.
D.Create an Amazon Kinesis Data Streams stream with a partition key of the instrument symbol, and have each consumer use the Kinesis Client Library.
AnswerD

Kinesis Data Streams retains records for up to 365 days and allows many consumers to read the same data independently, each tracking its own position. Using the instrument symbol as the partition key routes all events for a symbol to the same shard, which preserves per-symbol ordering. The Kinesis Client Library handles shard discovery, checkpointing, and load balancing across consumer instances.

Why this answer

Kinesis Data Streams is built for high-throughput, low-latency event ingestion where many applications read the same stream independently. Partitioning by instrument symbol guarantees that all events for a given symbol land on one shard and are read in order, and the Kinesis Client Library gives each consumer application its own checkpointed position without affecting the others.

Exam trap

The trap here is assuming that any fan-out mechanism preserves ordering, when standard SQS queues and many pub/sub designs offer no per-key ordering guarantee.

827
MCQmedium

A financial analytics company runs an Amazon RDS for MySQL database that supports a read-heavy web application. The primary DB instance is heavily loaded during business hours, and read replicas are already deployed and receiving traffic from the application. The team wants to reduce the load on the primary DB instance caused by read queries as much as possible, while keeping the application changes minimal. Which action should a solutions architect take?

A.Increase the size of the primary DB instance to a larger instance class.
B.Enable Multi-AZ deployment on the primary DB instance.
C.Configure the application to send read queries to the existing read replicas instead of the primary DB instance.
D.Create an Amazon ElastiCache for Redis cluster and cache all read results.
AnswerC

Read replicas are designed to serve read-only traffic using MySQL's asynchronous replication. Directing the application's read queries to the existing replicas removes that read load from the primary DB instance, which is the goal here. Because replicas are already deployed and receiving some traffic, extending application read routing to them is the minimal-change, correct optimization for reducing primary load.

Why this answer

Read replicas exist precisely to offload read-only query traffic from a primary DB instance. Since replicas are already deployed and receiving some traffic, routing the application's read queries to them reduces contention on the primary with minimal change. Multi-AZ is for availability, vertical scaling adds cost without offloading reads, and introducing a cache is a larger architectural change that does not use existing replica capacity.

Exam trap

The trap here is assuming Multi-AZ standbys can serve read traffic to offload the primary, when they are actually passive failover targets.

828
MCQeasy

A financial services company stores sensitive customer statements in an Amazon S3 bucket. The security team requires that all data be encrypted at rest using keys that the company manages and rotates on its own schedule. The company also needs an audit trail of every time a key was used to encrypt or decrypt data. Which solution meets these requirements?

A.Use S3 server-side encryption with AWS KMS customer managed keys (SSE-KMS).
B.Enable default encryption on the bucket using Amazon S3 managed keys and enable versioning.
C.Use S3 server-side encryption with Amazon S3 managed keys (SSE-S3).
D.Use S3 server-side encryption with customer-provided keys (SSE-C).
AnswerA

SSE-KMS with customer managed keys allows the company to create and rotate keys on its own schedule. AWS KMS records every use of the key in AWS CloudTrail, providing the required audit trail. This satisfies both the control over rotation and the need to log each cryptographic operation.

Why this answer

SSE-KMS with customer managed keys gives the company ownership of the key material lifecycle, including rotation, and integrates with AWS CloudTrail so every Encrypt and Decrypt call is logged. This combination directly addresses both the control and audit requirements. Other S3 encryption options either do not allow customer-managed rotation or do not produce a usable key usage audit trail.

Exam trap

The trap here is confusing encryption at rest with key usage auditing; only KMS-backed encryption with customer managed keys provides a CloudTrail record of each key operation.

829
MCQmedium

Based on the exhibit, which Route 53 configuration should be used so traffic automatically returns to the secondary Region only when the primary Region becomes unhealthy?

A.Use latency-based routing with both ALB records enabled.
B.Use failover routing with a primary alias record, a secondary alias record, and a Route 53 health check on the primary target.
C.Use geolocation routing so users are always sent to the closest Region.
D.Use a CNAME record that points to both ALBs so DNS can round-robin between Regions.
AnswerB

Failover routing is designed for this pattern: Route 53 returns the primary alias while the primary endpoint is healthy, and switches to the secondary alias when the primary health check fails. Alias records integrate cleanly with ALB targets, and the health check provides the signal that drives the failover decision.

Why this answer

Failover routing in Amazon Route 53 is designed for active-passive configurations. By creating a primary alias record pointing to the ALB in the primary Region and a secondary alias record pointing to the ALB in the secondary Region, and attaching a Route 53 health check to the primary target, traffic automatically fails over to the secondary Region only when the health check detects the primary as unhealthy. This meets the requirement of returning traffic to the secondary Region only upon primary failure.

Exam trap

The trap here is that candidates often confuse failover routing with latency-based or geolocation routing, assuming that 'closest' or 'fastest' automatically implies health awareness, but Route 53 health checks must be explicitly associated with failover records to trigger automatic traffic redirection.

How to eliminate wrong answers

Option A is wrong because latency-based routing directs users based on lowest latency, not health status, so it would not automatically fail over only when the primary is unhealthy; traffic could still be sent to an unhealthy primary if latency is low. Option C is wrong because geolocation routing sends users based on their geographic location, not the health of the endpoint, so it cannot automatically redirect traffic to the secondary Region when the primary becomes unhealthy. Option D is wrong because a CNAME record cannot point to multiple ALBs for round-robin; CNAME records can only point to a single DNS name, and DNS round-robin does not consider health checks, so traffic would still be sent to an unhealthy primary.

830
MCQeasy

A travel booking site uses EC2 instances behind an ALB. CPU is consistently high during peak traffic, and request latency rises. What should be configured?

A.A VPC endpoint for CloudWatch only
B.Auto Scaling policy based on an appropriate CloudWatch metric
C.S3 Object Lock
D.Disable health checks
AnswerB

An Auto Scaling policy driven by a suitable CloudWatch metric, such as average CPU utilisation or ALB request count per target, adds instances when demand rises and removes them afterwards. This directly addresses the sustained high CPU and latency during peak traffic.

Why this answer

An Auto Scaling policy based on a CloudWatch metric like CPUUtilization or request latency directly addresses the root cause: rising CPU and latency under peak traffic. By automatically adding EC2 instances when the metric breaches a threshold, the ALB can distribute load across more resources, reducing CPU per instance and improving response times. This is the standard AWS solution for dynamic scaling to maintain performance.

Exam trap

The trap here is that candidates may confuse monitoring (VPC endpoints) or data protection (S3 Object Lock) with scaling solutions, or think disabling health checks reduces overhead, when the correct approach is to scale horizontally based on load metrics.

How to eliminate wrong answers

Option A is wrong because a VPC endpoint for CloudWatch only enables private connectivity to CloudWatch without internet gateway, but does not add compute capacity or reduce CPU load or latency. Option C is wrong because S3 Object Lock prevents object deletion or overwrite for compliance, which is irrelevant to EC2 CPU and latency issues. Option D is wrong because disabling health checks would cause the ALB to route traffic to unhealthy instances, increasing failures and latency, not solving the performance problem.

831
MCQmedium

A video processing pipeline runs batch jobs that are safe to interrupt and restart. The jobs checkpoint progress to durable storage every few minutes, and the team can automatically resubmit from the last checkpoint. They want to minimize compute cost while accepting that capacity can be interrupted. Which launch configuration for the processing workers is the best cost-optimized choice?

A.Launch the worker nodes as Spot Instances, and configure the job resubmission logic to restart from checkpoints upon interruption.
B.Launch the worker nodes as On-Demand Instances with no interruption handling so the pipeline never needs resubmission.
C.Launch the worker nodes as Reserved Instances to guarantee capacity and reduce cost, ignoring interruptions.
D.Use Savings Plans and also set the job scheduler to never start new jobs unless previous jobs finish without interruption.
AnswerA

Spot provides significantly lower pricing than On-Demand for EC2 capacity. Because the workload is designed to tolerate interruption (checkpointing + resubmission from the last checkpoint), the team can safely accept Spot interruptions. Resubmission from durable checkpoints preserves correctness while still capturing the cost advantage of Spot.

Why this answer

Spot Instances offer significant cost savings (up to 90% compared to On-Demand) and are ideal for fault-tolerant, interruptible workloads. Since the pipeline checkpoints progress to durable storage and can automatically resume from the last checkpoint, using Spot Instances minimizes compute cost while accepting interruptions.

Exam trap

The trap here is that candidates may choose On-Demand or Reserved Instances because they assume interruptions are unacceptable, but the question explicitly states the workload is safe to interrupt and restart, making Spot Instances the correct cost-optimized choice.

How to eliminate wrong answers

Option B is wrong because On-Demand Instances are more expensive and provide no cost optimization benefit for a workload that can tolerate interruptions. Option C is wrong because Reserved Instances require a 1- or 3-year commitment and are not designed for workloads that can be interrupted; they also do not inherently handle interruption recovery. Option D is wrong because Savings Plans still incur costs for unused capacity if jobs are delayed, and the suggestion to never start new jobs unless previous jobs finish without interruption contradicts the goal of minimizing cost by accepting interruptions.

832
MCQhard

A batch analytics job currently uses two NAT gateways in each of three Availability Zones, but only one private subnet per AZ needs outbound internet access. What should the architect review first?

A.Replacing every NAT gateway with an internet gateway attached to private subnets
B.Whether one NAT gateway per AZ is sufficient for the required private subnets
C.Disabling route tables
D.Moving all workloads to public subnets
AnswerB

The right cost-optimization review here is to ask whether the batch analytics job truly needs a NAT gateway in every Availability Zone or whether one per AZ is sufficient. A NAT gateway is a zonal resource, and the standard high-availability pattern is to deploy one per AZ and route each AZ's private subnets to its local gateway; having two NAT gateways in the same AZ adds no resiliency because an AZ failure affects both, and the second gateway only doubles the hourly and per-GB costs. If the batch job is not required to be highly available or can tolerate an AZ outage, a single NAT gateway per AZ (or even one total) could be enough, which is exactly what should be evaluated before paying for four gateways.

Why this answer

The question asks what the architect should review first. Using two NAT gateways per Availability Zone (AZ) when only one private subnet per AZ needs outbound internet access is likely over-provisioned and costly. The architect should first verify if a single NAT gateway per AZ can handle the traffic load, as NAT gateways are highly available within an AZ and can support up to 45 Gbps of bandwidth.

This review directly addresses cost optimization without sacrificing functionality.

Exam trap

The trap here is that candidates may assume more NAT gateways always improve reliability, but the question emphasizes cost optimization, so the first review should be whether the existing number of gateways is necessary rather than immediately adding or removing resources.

How to eliminate wrong answers

Option A is wrong because replacing NAT gateways with an internet gateway attached to private subnets is technically invalid; internet gateways can only be attached to VPCs and provide outbound access only to resources with public IPs in public subnets, not private subnets. Option C is wrong because disabling route tables would break all network connectivity, not just outbound internet access, and is not a valid cost-optimization review step. Option D is wrong because moving all workloads to public subnets would expose them directly to the internet, violating security best practices and potentially incurring higher data transfer costs, and does not address the cost of NAT gateways.

833
MCQmedium

A web application runs on an Auto Scaling group (ASG) behind an Application Load Balancer (ALB). After a new release, instances begin failing ALB health checks with errors like 502 while the application is still starting up. CloudWatch shows that the ASG replaces the instances before they finish initializing, so traffic never reaches healthy targets. Which change most directly prevents premature replacement during startup so traffic can resume as soon as the instances are actually healthy?

A.Reduce the ALB health check timeout to 1 second so failures are detected faster.
B.Increase the Auto Scaling group health check grace period to cover application startup and initialization time.
C.Enable connection draining on the ALB target group but set deregistration delay to 0 seconds.
D.Switch the ALB target group health checks from HTTP to TCP so the application does not need to return HTTP 200.
AnswerB

The ASG health check grace period tells Auto Scaling to ignore failing health checks for a period after instance launch. This prevents newly launched instances from being replaced before the application has finished booting and can pass ALB health checks.

Why this answer

B is correct because the Auto Scaling group health check grace period allows instances a specified amount of time to initialize before the ASG starts checking their health status. By increasing this grace period to cover the application startup time, the ASG will not prematurely replace instances that are still initializing, allowing them to pass the ALB health checks and begin receiving traffic once they are actually healthy.

Exam trap

The trap here is that candidates often confuse the ALB health check timeout or interval with the ASG health check grace period, thinking that adjusting ALB settings will fix the premature replacement issue, when in fact the ASG grace period is the direct control for delaying health check evaluation during startup.

How to eliminate wrong answers

Option A is wrong because reducing the ALB health check timeout to 1 second would cause health checks to fail even faster, exacerbating the problem of premature instance replacement. Option C is wrong because connection draining controls how existing connections are closed during deregistration, not how quickly instances are replaced during startup; setting deregistration delay to 0 seconds would abruptly terminate active connections, causing user disruption. Option D is wrong because switching to TCP health checks would bypass the application layer, allowing the ALB to consider an instance healthy even if the application is not fully initialized, which could lead to serving 502 errors to users.

834
MCQeasy

A company runs a stateless web API on Amazon EC2 behind an Application Load Balancer. The team notices that during business hours, the ALB starts queueing requests and the average request latency rises. They want to scale out quickly and reliably based on demand, not CPU alone. Which Auto Scaling approach best matches this requirement?

A.Use a fixed-size Auto Scaling group and increase capacity manually once per hour.
B.Use target tracking scaling based on ALB request count per target.
C.Scale based only on EC2 instance memory utilization, regardless of load.
D.Use step scaling with a single threshold on average network-in bytes.
AnswerB

Target tracking scaling with ALB request count per target directly measures actual demand by dividing incoming requests by the number of healthy EC2 instances. The policy automatically adjusts capacity to keep the average near a target value, and because the ALB metric updates in near real time, it can add instances within minutes when a traffic spike begins. This ties scaling directly to the user-facing load that causes queuing and latency, making it the most responsive and precise option.

Why this answer

Target tracking scaling based on ALB request count per target directly aligns with the requirement to scale out based on demand (request queuing and latency) rather than CPU alone. This policy automatically adjusts the Auto Scaling group size to maintain a target value for the average number of requests per instance, which is a more reliable indicator of load for a stateless web API than CPU utilization.

Exam trap

The trap here is that candidates often assume CPU utilization is the best metric for all scaling scenarios, but for a stateless web API behind an ALB, request count per target is a more direct and reliable indicator of demand and latency issues.

Why the other options are wrong

A

Manual scaling once per hour cannot respond quickly to sudden demand spikes during business hours, leading to request queuing and increased latency.

C

Scaling based on EC2 instance memory utilization is not appropriate for a stateless web API where the bottleneck is request queuing and latency, not memory. Memory utilization may not correlate with demand, leading to under- or over-scaling.

D

Network-in bytes is not a reliable indicator of request queuing or latency for a stateless web API; it can spike due to large payloads without corresponding load, and step scaling with a single threshold lacks the smooth, demand-responsive behavior needed to prevent queueing.

When would these options actually be correct?

A

For a stateless application with predictable, gradual traffic changes where cost control is critical and automated scaling is not desired, a fixed-size group with manual adjustments might be acceptable.

C

This option would be correct for an application that is memory-bound, such as an in-memory cache or a data processing job where high memory usage indicates the need for more instances, and CPU or request metrics are not the primary drivers.

D

A company runs a data ingestion service that receives large file uploads, and scaling should trigger when network throughput exceeds a critical threshold to avoid packet loss. Step scaling with a single threshold on average network-in bytes would be appropriate to add capacity quickly when network input spikes.

Why candidates pick the wrong answer

A

Candidates may think manual scaling is simpler and more controllable, underestimating the need for rapid, automated scaling in response to real-time demand.

C

Candidates might think memory utilization is a good proxy for load, especially if they have experience with memory-intensive applications, but for a stateless web API, request count is a more direct indicator of demand.

D

Candidates may think network-in bytes reflects incoming request load, but it ignores request count and latency, and step scaling seems simpler to configure than target tracking, leading them to overlook the need for a metric that directly measures demand.

835
MCQmedium

A SaaS company runs a production API on an EC2 Auto Scaling group with steady demand 24/7. The team uses multiple instance types over time (they switch types during tuning) but the overall compute hours are stable. They want a cost reduction without committing to a specific instance type or size. Which AWS pricing option best meets the requirement?

A.Buy EC2 Spot Instances for the Auto Scaling group to maximize savings
B.Purchase a Compute Savings Plan for the region and commit to a dollar-per-hour amount
C.Purchase Reserved Instances that are limited to a single specific instance type in the Auto Scaling group
D.Use on-demand only, and rely on Auto Scaling to reduce cost during low utilization
AnswerB

A Compute Savings Plan lets you commit to a specific dollar-per-hour amount for a one- or three-year term in a given region, and the discount automatically applies to any EC2 instance family or size in that region. This is ideal for an Auto Scaling group with steady 24/7 API traffic because it captures predictable usage while preserving the flexibility to scale or change instance families without renegotiating the commitment. Usage above the committed amount simply runs at normal on-demand rates, so you still receive the lower rate on the bulk of your steady baseline.

Why this answer

B is correct because a Compute Savings Plan provides the flexibility to change instance types, sizes, and even compute services (e.g., EC2, Fargate, Lambda) within a region while still receiving discounted rates (up to 66% vs. on-demand). This matches the requirement of reducing costs without committing to a specific instance type or size, as the plan is based on a dollar-per-hour commitment rather than instance family or tenancy.

Exam trap

The trap here is that candidates often confuse Compute Savings Plans with Reserved Instances, assuming that any savings plan requires a specific instance type, but Compute Savings Plans offer full flexibility across instance families and sizes within a region.

How to eliminate wrong answers

Option A is wrong because Spot Instances can be interrupted with a 2-minute warning, making them unsuitable for a production API with steady demand 24/7 where availability and reliability are critical. Option C is wrong because Reserved Instances are tied to a specific instance type (e.g., m5.large) and tenancy, which contradicts the requirement to avoid committing to a specific instance type or size. Option D is wrong because relying solely on on-demand instances with Auto Scaling does not reduce cost; Auto Scaling only adjusts capacity based on demand, but on-demand pricing is the highest, so no cost savings are achieved.

836
MCQmedium

A company hosts a financial reporting platform on EC2. Administrators must connect without opening SSH or RDP ports to the internet. What should the architect use?

A.A public Elastic IP address on each instance
B.A bastion host with SSH open to 0.0.0.0/0
C.AWS Systems Manager Session Manager with the required instance role
D.An internet gateway attached to the private subnet
AnswerC

AWS Systems Manager Session Manager uses the SSM agent and an IAM instance role to provide secure shell access without needing inbound ports such as 22 or 3389. It authenticates through the AWS control plane, encrypts all session traffic, and can record sessions in CloudTrail and S3 for auditing. This provides the required audited, secure administrative access while keeping instances in a private subnet.

Why this answer

AWS Systems Manager Session Manager allows secure shell access to EC2 instances without opening inbound ports (SSH 22 or RDP 3389) to the internet. It uses the AWS Systems Manager agent on the instance, combined with an IAM instance role that grants permissions to communicate with the Systems Manager API, establishing a bidirectional tunnel over HTTPS (port 443). This satisfies the requirement of no public-facing SSH or RDP ports while enabling administrative connectivity.

Exam trap

The trap here is that candidates often default to a bastion host (Option B) as the traditional solution, but the question explicitly prohibits opening SSH or RDP ports to the internet, and a bastion host still requires those ports open (even if restricted to a CIDR), which fails the requirement; Session Manager avoids any inbound port exposure entirely.

How to eliminate wrong answers

Option A is wrong because assigning a public Elastic IP address to each instance would expose them directly to the internet, requiring open SSH or RDP ports to connect, which violates the requirement. Option B is wrong because a bastion host with SSH open to 0.0.0.0/0 exposes the bastion itself to the entire internet, creating a single point of attack and still requiring open SSH ports, which does not meet the 'without opening SSH or RDP ports to the internet' constraint. Option D is wrong because an internet gateway attached to a private subnet does not provide administrative connectivity; it enables outbound internet access for instances in public subnets, not inbound management access without open ports.

837
Multi-Selectmedium

A company is designing a multi-Region disaster recovery (DR) strategy for a stateless web application running on Amazon EC2 instances behind an Application Load Balancer (ALB). The application uses an Amazon RDS for MySQL database as its data store. The architecture must provide rapid failover with the lowest possible Recovery Point Objective (RPO) and Recovery Time Objective (RTO). Which of the following design choices will help achieve these objectives? (Choose four.)

Select 4 answers
.Configure an active-passive failover strategy by deploying the application stack in two AWS Regions and using Amazon Route 53 health checks with a failover routing policy.
.Set up Amazon RDS Multi-AZ deployment to enable automatic failover to a standby replica in a different Availability Zone within the primary Region.
.Use Amazon RDS cross-Region read replicas with automatic failover to promote a read replica to a primary instance in the secondary Region.
.Deploy the application and ALB in an active-active configuration across two AWS Regions using Amazon Route 53 latency-based routing.
.Store static assets and application state in Amazon S3 with cross-Region replication enabled, and serve them via Amazon CloudFront.
.Use an Amazon RDS for MySQL single-AZ deployment in the primary Region and take daily snapshots copied to the secondary Region.

Why this answer

An active-passive failover strategy with Route 53 failover routing policy is correct because it provides rapid failover by directing traffic to the secondary Region only when health checks fail in the primary, minimizing RTO. Cross-Region read replicas with automatic failover are correct because they allow promoting a read replica to a primary in the secondary Region with low RPO (typically seconds) and automated failover, reducing RTO. Active-active configuration with latency-based routing is correct because it distributes traffic across both Regions, enabling immediate failover without DNS propagation delays, achieving very low RTO.

Storing static assets and application state in S3 with cross-Region replication and CloudFront is correct because it ensures data durability and low-latency access, supporting rapid recovery with minimal RPO.

Exam trap

The trap here is that candidates often confuse Multi-AZ (single-Region high availability) with cross-Region DR, or they assume daily snapshots provide adequate RPO for a DR strategy requiring the lowest possible RPO and RTO.

838
Multi-Selectmedium

A CPU-bound batch rendering service runs on EC2. The application is Linux-based, compatible with ARM64, and the team wants the best throughput per dollar without changing the workload's architecture. Which two instance-family choices should the team consider first? Select two.

Select 2 answers
A.A compute-optimized family, because it is designed for workloads that spend most of their time on CPU.
B.A Graviton-based family, because compatible ARM instances often provide better price performance for many compute workloads.
C.A memory-optimized family, because extra RAM always increases compute throughput.
D.A storage-optimized family, because local storage bandwidth is the main factor for rendering performance.
E.A burstable family, because CPU credits make sustained rendering faster during long runs.
AnswersA, B

Compute-optimized families provide the highest ratio of vCPUs to memory and are engineered for workloads that need sustained, high-throughput CPU processing. These instances feature high-performance processors and low per-core latency, making them ideal for batch rendering where the application spends most of its time executing compute instructions. By selecting a compute-optimized family, you ensure the rendering service gets the maximum CPU resources per dollar without paying for unnecessary memory or storage.

Why this answer

Compute-optimized families (e.g., C5, C6g) are designed for workloads that spend most of their time on CPU, such as batch rendering. Option B is correct because Graviton-based instances (e.g., C6g, M6g) use ARM64 architecture, which is compatible with the workload and often delivers better price-performance for compute-intensive tasks, maximizing throughput per dollar without architectural changes.

Exam trap

The trap here is that candidates may confuse 'CPU-bound' with 'memory-bound' or 'storage-bound,' leading them to select memory-optimized or storage-optimized families, or they may mistakenly think burstable instances can sustain high CPU performance over long periods.

839
MCQeasy

A startup runs a static marketing website on Amazon S3 and wants to serve it to users worldwide with low latency. The site consists of HTML, CSS, JavaScript, and images stored in a single S3 bucket in the us-east-1 Region. The team wants to minimize cost while improving global performance. Which solution should the team implement?

A.Create an Amazon CloudFront distribution with the S3 bucket as the origin and use the default cache behavior.
B.Enable S3 Transfer Acceleration on the bucket and update the website URLs to use the accelerated endpoint.
C.Enable S3 Cross-Region Replication to replicate the bucket to multiple Regions and direct users to the closest bucket using latency-based routing.
D.Move the website content to an Amazon EBS volume attached to an EC2 instance in each Region and serve it with a web server.
AnswerA

CloudFront caches content at edge locations worldwide, reducing latency for global users by serving requests from the nearest edge. Using the S3 bucket as the origin with default cache behavior is a low-cost, standard pattern for static sites. It requires minimal configuration and scales automatically, making it the best fit for the startup's goals.

Why this answer

CloudFront is the standard AWS service for delivering static content globally with low latency. It caches objects at edge locations, reducing the number of requests that reach the S3 origin and improving performance for users everywhere. Cross-Region Replication, EC2-hosted content, and Transfer Acceleration either add cost and complexity or do not provide edge caching, making them less suitable for this scenario.

Exam trap

The trap here is confusing S3 Transfer Acceleration, which speeds up data transfer, with CloudFront, which caches content at edge locations for low-latency delivery.

840
MCQhard

A Lambda-based travel booking site has unpredictable traffic spikes and users see latency caused by cold starts. The function must respond consistently during expected campaign windows. What should be configured? The architecture review board prefers a managed AWS-native control.

A.Provisioned concurrency during campaign windows
B.A larger deployment package
C.CloudTrail data events
D.Reserved concurrency only
AnswerA

Provisioned concurrency during campaign windows is the correct solution because it initializes the Lambda runtime and application code in advance, removing the cold-start penalty for the first request in each concurrently processed execution environment. By setting a provisioned concurrency level prior to the campaign, the booking site can handle unpredictable spikes with consistently low latency, and the configuration can be scaled or scheduled to align with expected traffic patterns.

Why this answer

Provisioned concurrency initializes a specified number of execution environments in advance, eliminating cold starts for those instances. During campaign windows, this ensures consistent latency by keeping the function warm and ready to handle spikes without the delay of initializing new environments. It is a managed AWS-native control that directly addresses the unpredictable traffic pattern described.

Exam trap

The trap here is confusing reserved concurrency (which limits scaling but does not prevent cold starts) with provisioned concurrency (which pre-warms instances to eliminate cold starts), leading candidates to choose reserved concurrency as a simpler but incorrect solution.

How to eliminate wrong answers

Option B is wrong because a larger deployment package increases cold start time due to longer download and initialization overhead, making latency worse, not better. Option C is wrong because CloudTrail data events record API activity for auditing and governance, not for managing function initialization or latency. Option D is wrong because reserved concurrency only guarantees a maximum number of concurrent executions for a function, preventing it from using all available concurrency, but does not pre-warm instances; cold starts still occur for new invocations.

841
MCQhard

Based on the exhibit, a single EC2 instance hosts a latency-sensitive cache that performs sustained random reads and writes to persistent block storage. The current EBS volume is a general-purpose SSD, but BurstBalance is repeatedly depleted and p95 I/O latency has risen above 20 ms. The workload needs more than 16,000 sustained IOPS. Which change is the best fix?

A.Move the data to Amazon S3 so the instance can read and write objects directly.
B.Replace the volume with an io2 EBS volume and provision the required IOPS.
C.Keep gp2 and increase the instance size to a compute-optimized family.
D.Enable Amazon EFS with bursting throughput mode for the cache data.
AnswerB

io2 is designed for mission-critical workloads that need sustained, predictable, low-latency random I/O. Unlike gp2, it does not depend on burst credits for performance. Provisioning the required IOPS directly addresses the exhausted BurstBalance and the sustained throughput requirement above 16,000 IOPS.

Why this answer

The workload requires more than 16,000 sustained IOPS with low latency, and the gp2 volume's burst credits are exhausted, causing high latency. An io2 Block Express or io2 volume can be provisioned with the exact IOPS needed (up to 256,000 IOPS) and provides consistent single-digit millisecond latency, making it the best fix for this latency-sensitive, sustained I/O workload.

Exam trap

The trap here is that candidates often assume increasing instance size (Option C) will improve EBS performance, but EBS IOPS and throughput are tied to the volume type and size, not the instance type (except for EBS-optimized bandwidth), so the gp2 burst credit exhaustion remains the root cause.

How to eliminate wrong answers

Option A is wrong because Amazon S3 is object storage accessed via HTTPS, not block storage, and introduces network latency and throughput limitations that are unsuitable for a latency-sensitive cache requiring sustained random reads/writes. Option C is wrong because increasing the instance size to a compute-optimized family does not change the gp2 volume's burst credit model; the volume will still deplete its burst balance and throttle to baseline IOPS (e.g., 160 IOPS per GB), failing to meet the >16,000 sustained IOPS requirement. Option D is wrong because Amazon EFS is a shared file system with NFS protocol overhead and its bursting throughput mode relies on burst credits that can be exhausted, leading to throttled throughput and higher latency, not suitable for sustained high IOPS block-level cache workloads.

842
Multi-Selecthard

Multiple EC2 instances in different Availability Zones need concurrent read/write access to the same shared files. The files are actively modified by several application servers, and low-latency metadata operations matter more than extremely high aggregate throughput. Which two changes should the team make? Select two.

Select 2 answers
A.Use Amazon EFS instead of EBS or S3 for the shared file system.
B.Create EFS mount targets in every Availability Zone that hosts application instances.
C.Use a single EBS Multi-Attach volume mounted read/write by all instances across AZs.
D.Store the files in S3 and mount them directly through the console as a shared network filesystem.
E.Place the files on instance store volumes so each server has faster local access.
AnswersA, B

Amazon EFS is the managed AWS file service built for shared POSIX-style file access from multiple instances. It supports concurrent read/write access from many EC2 hosts and is a better fit than EBS, which is attached to a single instance, or S3, which provides object storage rather than a native shared filesystem. For an application that expects standard filesystem semantics, EFS is the correct storage layer.

Why this answer

Amazon EFS provides a fully managed, POSIX-compliant, shared file system that can be mounted concurrently by multiple EC2 instances across different Availability Zones (AZs). It supports concurrent read/write access with strong consistency, and its metadata operations are optimized for low latency, making it ideal for workloads where many application servers actively modify the same files. EBS cannot be shared across AZs, and S3 lacks POSIX semantics and low-latency metadata operations.

Exam trap

The trap here is that candidates often confuse EBS Multi-Attach with a cross-AZ shared storage solution, but Multi-Attach is strictly limited to a single AZ and a small number of instances, while EFS is the only AWS shared file system that natively spans AZs with concurrent read/write access.

843
MCQhard

An EC2 instance in a private subnet must access an S3 bucket that contains regulated exports for a financial reporting platform. The security team requires access to be allowed only when traffic comes through a specific VPC endpoint. What should the architect add to the bucket policy? The design must avoid adding custom operational scripts.

A.A security group rule that allows HTTPS to S3
B.A condition that matches aws:RequestedRegion to the bucket Region
C.A deny statement for all IAM users except the EC2 role
D.A condition that matches aws:sourceVpce to the endpoint ID
AnswerD

The aws:sourceVpce condition key in an S3 bucket policy checks the VPC endpoint ID from which the request originated. When the EC2 instance sends traffic through the designated VPC endpoint, the request includes the endpoint's ID, allowing the condition to match and granting access. Any request not coming through that endpoint, even with valid IAM credentials, would fail the condition and be denied, thus forcing the traffic to traverse the private endpoint.

Why this answer

The bucket policy can use the `aws:sourceVpce` condition key to restrict access to requests originating from a specific VPC endpoint. This ensures that only traffic routed through that endpoint can access the S3 bucket, meeting the security team's requirement without custom scripts.

Exam trap

The trap here is that candidates may confuse security group rules (which control instance-level traffic) with bucket policy conditions (which control access to the S3 service), leading them to pick Option A instead of the correct VPC endpoint condition.

How to eliminate wrong answers

Option A is wrong because security group rules are applied at the instance level, not the bucket policy level, and cannot restrict access based on the VPC endpoint used. Option B is wrong because `aws:RequestedRegion` checks the region of the request, not the network path or endpoint, so it does not enforce that traffic comes through a specific VPC endpoint. Option C is wrong because denying all IAM users except the EC2 role does not control the network path; the EC2 role could still access S3 via the internet or a different endpoint, violating the requirement.

844
MCQmedium

A batch analytics job runs for several hours each night and can be interrupted and restarted. Which EC2 purchasing option should minimize cost?

A.On-Demand Instances only
B.Dedicated Hosts
C.Spot Instances
D.Provisioned IOPS volumes
AnswerC

Spot Instances offer spare EC2 compute capacity at up to a 90% discount compared to On-Demand, with the tradeoff that AWS can reclaim this capacity with a two-minute warning for other customers' workloads. Because the batch analytics job runs nightly, lasts several hours, and can be interrupted and resumed, it is exactly the kind of fault-tolerant, flexible workload that Spot was designed for. Using Spot Instances dramatically reduces compute costs while accommodating the job's tolerance for interruptions, making it the correct answer for a cost-optimized architecture.

Why this answer

Spot Instances are the correct choice because they offer significant cost savings (up to 90% compared to On-Demand) and are ideal for fault-tolerant, interruptible workloads like batch processing. Since the job can be interrupted and restarted, it can handle Spot Instance terminations gracefully, making this the most cost-effective option.

Exam trap

The trap here is that candidates may choose On-Demand Instances thinking they need guaranteed uptime, overlooking the fact that the workload is explicitly described as interruptible and restartable, which makes Spot Instances the optimal cost-saving choice.

How to eliminate wrong answers

Option A is wrong because On-Demand Instances provide no interruption but are priced higher, which is unnecessary for a workload that can tolerate interruptions. Option B is wrong because Dedicated Hosts are designed for licensing or compliance requirements and are billed per host, making them far more expensive and unsuitable for cost minimization. Option D is wrong because Provisioned IOPS volumes (EBS) relate to storage performance, not compute pricing, and do not address the cost of EC2 instances.

845
MCQeasy

A company serves mostly static images and JavaScript files from an origin in one AWS Region. They want to reduce origin load and improve global performance. Which change most directly increases cache-hit ratio for static assets while avoiding stale content?

A.Set Cache-Control headers on the origin to always be no-cache so clients revalidate frequently.
B.Use versioned file names (e.g., app.abc123.js) and configure a long TTL with appropriate revalidation behavior.
C.Disable query string forwarding so all URLs without query strings share one cached object even when content differs.
D.Forward all headers, including cookies, to maximize personalization in edge cached responses.
AnswerB

Versioned filenames let each asset be cached with a long TTL because a content change produces a new URL, so the cache-hit ratio rises for static assets while revalidation behaviour prevents stale content being served under an unchanged name.

Why this answer

Using versioned file names (e.g., app.abc123.js) allows you to set a long Cache-Control max-age TTL (e.g., one year) without risking stale content. When the file changes, the new version gets a new URL, so clients and edge caches immediately fetch the fresh object, maximizing cache hits for unchanged assets while avoiding stale content.

Exam trap

The trap here is that candidates often confuse 'no-cache' with 'no-store' or think that disabling query strings universally improves caching, but they fail to recognize that versioned filenames with long TTLs are the standard pattern for maximizing cache hits while ensuring content freshness.

Why the other options are wrong

A

Setting Cache-Control: no-cache forces clients to revalidate with the origin on every request, which increases origin load and defeats caching, directly contradicting the goal of reducing origin load and improving performance.

C

Disabling query string forwarding causes all URLs without query strings to be treated as identical, even if the underlying content differs (e.g., different versions of a file). This can serve stale or incorrect content, reducing cache-hit ratio for static assets that rely on query parameters for versioning.

D

Forwarding all headers, including cookies, reduces cache-hit ratio because each unique set of headers creates a separate cached object, defeating the purpose of caching static assets that don't vary by user.

When would these options actually be correct?

A

In a scenario where content changes frequently and users must always see the latest version (e.g., real-time stock prices or live scores), and origin load is not a concern, using no-cache ensures freshness while still allowing conditional revalidation.

C

In a scenario where query strings are used for tracking or analytics (e.g., ?utm_source=facebook) and do not affect the actual content served, disabling query string forwarding would increase cache-hit ratio by treating all variations as the same object, improving cache efficiency without serving incorrect content.

D

In a scenario where content must be personalized per user (e.g., a dashboard with user-specific data), forwarding all headers ensures each user receives their tailored response from the edge, and the question asks for maximizing personalization rather than cache-hit ratio.

Why candidates pick the wrong answer

A

Candidates may think no-cache still allows caching with revalidation, but they overlook that it requires a round-trip to the origin for every request, increasing load and latency.

C

Candidates may think that ignoring query strings always improves cache-hit ratio by consolidating requests, without realizing that query strings are often used for versioning or content differentiation, and ignoring them can cause stale or wrong content to be served.

D

Candidates may think that forwarding all headers ensures the edge delivers the most accurate content, not realizing that for static assets, this dramatically reduces cache efficiency.

846
MCQmedium

A batch analytics job runs for several hours each night and can be interrupted and restarted. Which EC2 purchasing option should minimize cost? The design must avoid adding custom operational scripts.

A.On-Demand Instances only
B.Dedicated Hosts
C.Spot Instances
D.Provisioned IOPS volumes
AnswerC

Spot Instances let you bid on spare EC2 capacity at steep discounts — often 60-90% off On-Demand prices. Since the batch analytics job runs for several hours each night and can be interrupted and resumed, it is exactly the type of fault-tolerant workload AWS designed Spot Instances for. You can use Spot with a checkpointing strategy to save intermediate results, making it the most cost-effective choice.

Why this answer

Spot Instances are ideal for fault-tolerant, interruptible batch workloads because they offer significant cost savings (up to 90% off On-Demand pricing) by using spare EC2 capacity. Since the job can be interrupted and restarted, it can handle Spot Instance reclaimations without requiring custom operational scripts—AWS handles the interruption notification and automatic instance termination, and the job's restart logic can be built into the application or orchestration layer (e.g., AWS Batch).

Exam trap

The trap here is that candidates may confuse Spot Instances with On-Demand Instances for cost savings, or incorrectly assume that Spot Instances require custom scripting to handle interruptions, when in fact AWS provides built-in mechanisms (e.g., lifecycle hooks, rebalance notifications) that can be leveraged without custom scripts.

How to eliminate wrong answers

Option A is wrong because On-Demand Instances provide no cost savings for interruptible workloads; they are priced at the standard rate and are intended for steady-state or unpredictable workloads that cannot tolerate interruptions. Option B is wrong because Dedicated Hosts are a physical server dedicated to your use, which is significantly more expensive and unnecessary for a batch job that can tolerate interruptions; they are used for licensing or compliance requirements, not cost optimization. Option D is wrong because Provisioned IOPS volumes (EBS) are a storage type, not an EC2 purchasing option; they affect storage performance and cost but do not address compute cost optimization for interruptible workloads.

847
MCQeasy

A team wants to run containerized services with AWS-managed orchestration and autoscaling. They do NOT require Kubernetes compatibility. Which AWS service choice is most appropriate to meet these goals?

A.Amazon EKS
B.Amazon ECS
C.An EC2 Auto Scaling group only
D.Amazon SQS as the compute layer
AnswerB

Amazon ECS is a native container orchestration service. You can run containers without Kubernetes, and ECS integrates with AWS-native autoscaling (for example, ECS Service Auto Scaling with targets such as CPU/memory or request-based metrics when applicable to the architecture).

Why this answer

Amazon ECS is the most appropriate choice because it provides AWS-managed container orchestration and autoscaling without requiring Kubernetes compatibility. ECS integrates natively with AWS services like Application Auto Scaling and CloudWatch to automatically scale container tasks based on metrics such as CPU or memory utilization, meeting the team's requirements directly.

Exam trap

The trap here is that candidates often confuse Amazon ECS with Amazon EKS, assuming that Kubernetes compatibility is required for container orchestration, but ECS provides a simpler, AWS-native alternative without Kubernetes overhead.

How to eliminate wrong answers

Option A is wrong because Amazon EKS is a managed Kubernetes service that requires Kubernetes compatibility, which the team explicitly does not need, adding unnecessary complexity and overhead. Option C is wrong because an EC2 Auto Scaling group only manages EC2 instances, not container orchestration or scheduling, so it cannot run containerized services directly without additional container management software. Option D is wrong because Amazon SQS is a message queuing service, not a compute layer; it cannot run containers or provide orchestration or autoscaling for containerized workloads.

848
MCQeasy

A startup runs a stateless image-resizing API on a fleet of EC2 instances behind an Application Load Balancer. The instances store uploaded source images on their own instance store volumes before processing. During a routine scale-in event, an instance was terminated and several in-flight uploads were lost. The architect must make the design resilient to instance loss without changing the API code. What should the architect do?

A.Store uploaded images in Amazon S3 and have instances read from and write to the bucket instead of local disk
B.Increase the Auto Scaling group's minimum capacity so that instances are rarely terminated
C.Enable detailed CloudWatch monitoring and create an alarm that notifies operators before scale-in occurs
D.Configure an EC2 Auto Scaling lifecycle hook to delay instance termination until uploads complete
AnswerA

Moving the uploaded source images to Amazon S3 decouples the data from any single EC2 instance, so terminating an instance no longer destroys in-flight uploads. S3 provides durable, highly available object storage that all instances can access concurrently. Because the API already treats instances as stateless workers, replacing local storage with S3 is the standard resilience pattern and requires no change to the fleet's scaling behavior.

Why this answer

The root cause is that uploaded images live on instance store volumes, which are ephemeral and destroyed when an instance terminates. Relocating the source images to Amazon S3 removes the dependency on any single instance and gives the fleet shared, durable storage, so scale-in no longer causes loss. The remaining options delay or observe termination or reduce its frequency, but none preserve the data itself.

Exam trap

The trap here is treating instance store as durable because it is physically attached to the instance; instance store data is lost on stop, terminate, or host failure, so it cannot back resilient state.

849
MCQeasy

A retail API uses EC2 instances behind an ALB. CPU is consistently high during peak traffic, and request latency rises. What should be configured? The design must avoid adding custom operational scripts.

A.Auto Scaling policy based on an appropriate CloudWatch metric
B.S3 Object Lock
C.A VPC endpoint for CloudWatch only
D.Disable health checks
AnswerA

An Auto Scaling policy that uses a CloudWatch metric such as Average CPUUtilization immediately addresses the symptom: when CPU load rises beyond a threshold, it launches additional EC2 instances to distribute incoming ALB traffic, and when utilization falls it terminates excess instances. This is the standard, automated horizontal-scaling mechanism for a retail API and directly matches the need to accommodate fluctuating CPU load without manual intervention.

Why this answer

An Auto Scaling policy based on a CloudWatch metric like CPUUtilization or ALB TargetResponseTime can dynamically add or remove EC2 instances to match demand. This directly addresses the high CPU and rising latency during peak traffic without requiring custom scripts, as the scaling actions are fully managed by AWS. The ALB distributes traffic across the scaled instances, reducing per-instance load and improving response times.

Exam trap

The trap here is that candidates may confuse VPC endpoints (which enable private connectivity) with actual scaling mechanisms, or assume that disabling health checks is a quick fix for latency, when in fact it degrades reliability and does not address the underlying capacity issue.

How to eliminate wrong answers

Option B is wrong because S3 Object Lock is a data protection feature for S3 objects (preventing deletion/overwrite) and has no role in scaling compute resources or reducing latency for an API behind an ALB. Option C is wrong because a VPC endpoint for CloudWatch only enables private connectivity to CloudWatch APIs (e.g., for publishing metrics or logs) but does not automatically trigger scaling or resolve CPU/latency issues; scaling still requires an Auto Scaling policy. Option D is wrong because disabling health checks would cause the ALB to route traffic to unhealthy instances, worsening latency and availability, and it does not address the root cause of high CPU.

850
Multi-Selecthard

A logistics company runs workloads in a VPC with private subnets that have no internet gateway route. Instances in these subnets must retrieve secrets from AWS Secrets Manager and download patches from an Amazon S3 bucket owned by the company. The security team requires that this traffic never traverse the public internet. (Choose two.)

Select 2 answers
A.Create a VPC peering connection between the private subnets' VPC and the S3 service VPC in the same region.
B.Create a gateway VPC endpoint for Amazon S3 and associate it with the private subnets' route tables.
C.Deploy a NAT gateway in a public subnet and route the private subnets' traffic to it.
D.Attach an internet gateway to the VPC and add a route for 0.0.0.0/0 to the private subnets' route tables.
E.Create an interface VPC endpoint for Secrets Manager in the VPC and associate it with the private subnets.
AnswersB, E

A gateway endpoint for S3 adds a prefix list route to the subnet route tables that directs S3-bound traffic to the endpoint instead of an internet gateway or NAT device. Because the private subnets have no internet route, the gateway endpoint is what allows patch downloads from the company's bucket to succeed while keeping the traffic on the AWS network.

Why this answer

Private connectivity to AWS services uses VPC endpoints. Secrets Manager requires an interface endpoint because it is accessed over its API through PrivateLink, while S3 supports a gateway endpoint that adds a route to the subnet route tables. Together these endpoints keep both secret retrieval and patch downloads inside the AWS network without an internet gateway or NAT gateway.

Exam trap

The trap here is assuming that a NAT gateway keeps traffic private because it hides instance IPs, when in fact NAT gateway traffic still egresses to the public internet.

851
MCQmedium

A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The design must avoid adding custom operational scripts.

A.Lambda reserved concurrency set to zero
B.A Lambda dead-letter queue or failure destination
C.A larger deployment package
D.CloudFront error pages
AnswerB

A dead-letter queue or failure destination captures invocation records after Lambda exhausts its asynchronous retries, retaining failed events for later investigation. This is native Lambda configuration, so no custom operational scripts are needed, satisfying the requirement to retain failures without extra tooling.

Why this answer

A Lambda dead-letter queue (DLQ) or failure destination allows you to capture events that have exhausted all retry attempts from an asynchronous invocation. When the Lambda function fails after the maximum retries (default 3), the event is sent to the configured SQS queue or SNS topic for later investigation, without requiring custom scripts or manual polling.

Exam trap

The trap here is that candidates may confuse Lambda's DLQ/failure destination with other error-handling mechanisms like SQS redrive policies or CloudFront custom error pages, which serve different purposes and operate at different layers of the architecture.

How to eliminate wrong answers

Option A is wrong because setting reserved concurrency to zero would completely disable the Lambda function, preventing any invocations and thus failing to process or retain any events. Option C is wrong because a larger deployment package does not affect error handling or event retention; it only increases cold start latency and deployment size. Option D is wrong because CloudFront error pages are for HTTP-level errors from a web distribution, not for capturing asynchronous Lambda invocation failures or dead-letter events.

852
MCQmedium

A web application runs in private subnets with no NAT gateway. It needs to retrieve credentials from AWS Secrets Manager at runtime. After a recent network hardening change, the application logs timeout errors when calling Secrets Manager. Which change will most directly enable private connectivity to Secrets Manager while keeping the subnets NAT-free?

A.Create an interface VPC endpoint (AWS PrivateLink) for the Secrets Manager service and update the security group rules to allow HTTPS from the application subnets.
B.Add a public DNS entry in the instance /etc/hosts pointing Secrets Manager to the instance’s private IP so requests do not leave the VPC.
C.Attach an internet gateway to the private route table so that Secrets Manager traffic can reach public endpoints without NAT.
D.Enable S3 VPC endpoint and store the secrets in an S3 bucket instead of Secrets Manager, then retrieve them using S3 gateway endpoints.
AnswerA

An interface VPC endpoint (AWS PrivateLink) creates private ENIs in the subnets, routing Secrets Manager HTTPS traffic over the AWS network without internet or NAT. Security group rules permitting HTTPS from the application subnets complete the private path.

Why this answer

An interface VPC endpoint (AWS PrivateLink) for Secrets Manager creates a private, direct connection to the service within the VPC, using Elastic Network Interfaces (ENIs) in the subnets. This allows the application to reach Secrets Manager over HTTPS without traversing the internet, a NAT gateway, or an internet gateway, directly resolving the timeout errors caused by the network hardening change that removed public internet access.

Exam trap

The trap here is that candidates might think a NAT gateway or internet gateway is required for any AWS service access, overlooking that AWS PrivateLink interface endpoints can provide private, direct connectivity to services like Secrets Manager without any public internet exposure.

Why the other options are wrong

B

Modifying /etc/hosts on an instance does not create a private network path; traffic still routes through the internet unless a private connection exists. Without a NAT gateway or VPC endpoint, the instance cannot reach the public Secrets Manager endpoint, so the change does not resolve the timeout.

C

Attaching an internet gateway to a private route table would expose the private subnets to the internet, violating the requirement to keep subnets NAT-free and private, and it does not provide private connectivity to Secrets Manager.

D

This option suggests using S3 instead of Secrets Manager, but the question explicitly requires retrieving credentials from AWS Secrets Manager. Changing the service is not a direct solution to enable private connectivity to Secrets Manager.

When would these options actually be correct?

B

If the question described a scenario where the application needs to resolve a custom domain name to a private IP within the VPC (e.g., for a database or internal service) and the VPC already has a private network path (like a VPN or Direct Connect) to that IP, then adding a hosts entry would be a quick fix to avoid DNS resolution issues.

C

In a scenario where a public subnet's route table needs to allow direct outbound internet access for instances with public IPs, attaching an internet gateway to that route table is correct. For example, a web server in a public subnet that must reach public endpoints without NAT.

D

In a scenario where an application needs to retrieve configuration data or secrets stored in S3, and the subnets have no NAT gateway, an S3 VPC endpoint (gateway endpoint) would provide private connectivity to S3 without requiring a NAT or internet gateway.

Why candidates pick the wrong answer

B

Candidates may think that overriding DNS resolution with a private IP keeps traffic within the VPC, but they overlook that the underlying network path still requires a route to the destination, which is missing without a VPC endpoint or NAT.

C

Candidates may think that an internet gateway provides direct internet access without NAT, but they overlook that it must be attached to public subnets, not private ones, and that it does not create a private connection to AWS services.

D

Candidates may think that using S3 with a VPC endpoint is a valid workaround to avoid NAT, but they overlook the requirement to use Secrets Manager specifically, not S3.

853
Multi-Selecthard

A team is designing a new workload that runs on Amazon EC2 instances in a private subnet. The instances must read and write objects in an Amazon S3 bucket in the same account, and security policy forbids long-term access keys on the instances. The team wants to grant least-privilege access and ensure the instances can reach S3 without traversing the public internet. (Choose two.)

Select 2 answers
A.Create a NAT gateway in the public subnet and route the private subnet's default route to it so the instances can reach the S3 public endpoint.
B.Store an IAM user access key and secret access key in AWS Systems Manager Parameter Store as a SecureString, and have the application retrieve them at startup.
C.Attach an IAM role to the instances using an instance profile, and grant the role only the required s3:GetObject and s3:PutObject permissions on the specific bucket.
D.Configure an S3 bucket policy that allows the s3:GetObject and s3:PutObject actions for the account root user, and let the instances inherit those permissions.
E.Create an S3 gateway endpoint and associate it with the private subnet's route table so that traffic to S3 uses the endpoint instead of the internet.
AnswersC, E

An instance profile delivers temporary credentials through the instance metadata service, so no long-term access keys are stored on the instances. Scoping the role policy to the exact bucket and the two needed actions enforces least privilege. This directly addresses both the credential and the permission requirements for the workload.

Why this answer

An IAM role attached through an instance profile supplies temporary credentials without long-term keys, and an S3 gateway endpoint keeps bucket traffic on the AWS network. Together they satisfy the credential, least-privilege, and private-connectivity requirements. Static keys in Parameter Store and NAT-based egress each violate one of the stated constraints.

Exam trap

The trap here is assuming that storing keys in Parameter Store SecureString makes them acceptable, when the policy forbids long-term access keys on the instances entirely.

854
MCQmedium

A company runs a customer portal on an Amazon Aurora PostgreSQL cluster. The application currently connects directly to the writer instance endpoint and keeps long-lived connections open. During a maintenance failover, writes fail until clients are restarted. The team wants the application to reconnect to the correct Aurora endpoint automatically and reduce user-visible write interruptions. Which change is most likely to achieve this?

A.Use the Aurora cluster endpoint for write traffic, use the reader endpoint for read-only traffic, and implement connection retry or reconnect logic on failover.
B.Keep using the original writer instance endpoint so the database host name never changes during failover.
C.Convert the Aurora cluster to Single-AZ so there is only one database node to connect to.
D.Place Route 53 in front of the database and manually update DNS records whenever failover occurs.
AnswerA

The cluster endpoint always resolves to the current writer, so reconnecting after failover targets the promoted instance rather than a stale writer address; retry logic lets long-lived connections re-establish without restarting clients, reducing write interruption.

Why this answer

The Aurora cluster endpoint automatically points to the current writer instance, so using it for write traffic ensures that after a failover, new writes are directed to the new writer without needing to change the connection string. Implementing connection retry or reconnect logic in the application is essential because the existing long-lived connections will be broken during failover; the application must detect the failure and re-establish connections to the cluster endpoint to resume writes seamlessly.

Exam trap

The trap here is that candidates assume the writer instance endpoint remains constant during failover (Option B), but in Aurora, the writer instance endpoint changes because it is tied to the specific DB instance, not the cluster.

Why the other options are wrong

B

The writer instance endpoint points to a specific Aurora node, which changes during failover. Keeping it does not automatically redirect traffic to the new writer, so writes still fail until clients are restarted.

C

Converting to Single-AZ removes the standby replica, eliminating high availability. During a failover, there is no standby to promote, causing longer downtime and potential data loss, which contradicts the goal of reducing write interruptions.

D

Manually updating Route 53 DNS records during failover is not automated and would still cause write interruptions until the manual update is completed, failing to meet the requirement of automatic reconnection and reduced downtime.

When would these options actually be correct?

B

In a scenario where the database is a standalone RDS instance (not Aurora) and the application uses a CNAME pointing to the instance endpoint, the endpoint remains the same after failover if Multi-AZ is enabled, so no reconnect logic is needed.

C

An exam scenario where cost reduction is the primary goal and the application can tolerate downtime (e.g., a development or test environment) would make Single-AZ correct. The question would explicitly state that high availability is not required.

D

In a scenario where a company needs to redirect traffic to a standby database in a different region after a disaster, and they have a script or automation to update Route 53 records, this could be a valid approach for manual failover control.

Why candidates pick the wrong answer

B

Candidates may assume the writer endpoint is static and failover is transparent, not realizing Aurora's writer endpoint changes to a different physical node after failover.

C

Candidates may think that fewer nodes means simpler failover, overlooking that Aurora's Multi-AZ failover is automatic and faster. They might assume Single-AZ avoids failover issues entirely, not realizing it removes redundancy.

D

Candidates may think DNS-based routing provides a simple way to change endpoints without modifying application code, overlooking the need for automation and the fact that Aurora already provides cluster endpoints for this purpose.

855
MCQmedium

A media processing pipeline runs batch jobs on EC2. The jobs can tolerate interruptions because they checkpoint progress to durable storage and can restart. The total workload is variable week-to-week, and there is no need to guarantee capacity at specific times. To reduce compute cost while maintaining correctness, what EC2 purchase option and approach is the best fit?

A.Use EC2 Spot Instances with interruption handling and restart from checkpoints.
B.Use All Upfront Reserved Instances sized for the average weekly workload to minimize cost.
C.Use On-Demand Instances and scale only during business hours to reduce idle time.
D.Use Savings Plans with a fixed hourly commitment to ensure capacity for the entire year.
AnswerA

Spot capacity is typically the lowest-cost EC2 option and can be reclaimed by AWS with interruption notices. Because the workload is explicitly restartable and checkpoints to durable storage, interruptions do not break correctness. Since there is no requirement to reserve capacity, the variable workload aligns well with Spot’s spare-capacity model.

Why this answer

Spot Instances offer up to 90% cost savings compared to On-Demand and are ideal for fault-tolerant, stateless workloads that can checkpoint progress to durable storage. Since the batch jobs can tolerate interruptions and restart from checkpoints, Spot Instances provide the lowest compute cost while maintaining correctness. No other purchase option achieves the same level of cost reduction for this variable, interruption-tolerant workload.

Exam trap

The trap here is that candidates often choose Reserved Instances or Savings Plans thinking they always provide the best cost savings, but they fail to recognize that Spot Instances are significantly cheaper and perfectly suited for fault-tolerant, checkpointed batch workloads that do not require guaranteed capacity.

How to eliminate wrong answers

Option B is wrong because All Upfront Reserved Instances require a 1- or 3-year commitment and are sized for a fixed capacity, which does not match the variable week-to-week workload and would lead to over-provisioning or under-utilization, increasing cost. Option C is wrong because On-Demand Instances are the most expensive per-hour option and scaling only during business hours ignores the fact that the workload can run at any time; this approach does not minimize cost compared to Spot. Option D is wrong because Savings Plans with a fixed hourly commitment lock in a baseline spend and do not provide the deep discounts of Spot Instances; they also guarantee capacity only up to the committed amount, which is unnecessary for a workload that does not need guaranteed capacity.

856
MCQmedium

A web application for a IoT ingestion API is behind an Application Load Balancer. The application must be protected from common SQL injection and cross-site scripting attacks with minimum operational overhead. What should the architect deploy?

A.AWS WAF associated with the Application Load Balancer
B.Network ACLs on the public subnets
C.Security groups on the application instances
D.AWS Shield Advanced only
AnswerA

AWS WAF, when attached to an Application Load Balancer, acts as a web application firewall that inspects each HTTP(S) request at the application layer (Layer 7). It uses managed rule groups (e.g., AWS Managed Rules for SQL injection and XSS) and custom rules to filter and block malicious traffic before it reaches the backend instances. This is the appropriate service for detecting and mitigating SQL injection and cross-site scripting because it can parse HTTP headers, body, and URI patterns, and it integrates natively with ALB rules to allow, block, or count matching requests.

Why this answer

AWS WAF is a web application firewall that integrates directly with an Application Load Balancer to filter and monitor HTTP/HTTPS requests. It provides managed rules specifically designed to block common attack patterns like SQL injection and cross-site scripting (XSS) with minimal operational overhead, as AWS manages the rule updates and scaling. This makes it the ideal choice for protecting the IoT ingestion API without requiring custom code or manual configuration.

Exam trap

The trap here is that candidates often confuse network-layer controls (like NACLs or security groups) with application-layer protection, assuming that blocking ports or IPs is sufficient to prevent SQL injection and XSS, when in fact these attacks require deep packet inspection of HTTP content.

How to eliminate wrong answers

Option B is wrong because Network ACLs operate at the subnet level and provide stateless IP/port filtering; they cannot inspect application-layer payloads to detect SQL injection or XSS patterns. Option C is wrong because security groups act as stateful firewalls at the instance level, filtering traffic based on IP addresses and ports, but they lack the ability to parse HTTP request bodies or headers for malicious content. Option D is wrong because AWS Shield Advanced provides DDoS protection and enhanced monitoring, but it does not include web application firewall capabilities to block SQL injection or XSS attacks.

857
MCQmedium

A healthcare company runs a three-tier web application on AWS. The application tier consists of EC2 instances in an Auto Scaling group behind an Application Load Balancer. The security team must ensure that the application instances accept traffic only from the load balancer and that no instance can be reached directly from the internet. The instances are in private subnets and have a security group attached. What should a solutions architect do to meet these requirements?

A.Configure the instance security group to allow inbound traffic on the application port only from the load balancer's security group.
B.Configure the instance security group to allow inbound traffic on the application port from the VPC CIDR range.
C.Create a network ACL that denies all inbound traffic except from the load balancer's subnet CIDR range.
D.Attach an Elastic IP address to each instance and allow inbound traffic only from the load balancer's public IP addresses.
AnswerA

Referencing the load balancer's security group as the source in the instance security group's inbound rule allows only traffic that originates from the load balancer. Because security group references are evaluated on the source's security group membership, no IP ranges need to be maintained. This satisfies the requirement that instances accept traffic only from the ALB and cannot be reached directly from the internet.

Why this answer

The most secure and operationally simple way to restrict instance access to the load balancer is to reference the load balancer's security group in the instance security group's inbound rule. This creates a logical relationship rather than an IP-based one, so it remains correct as the ALB scales and its node IPs change. It also prevents direct internet access because private instances have no public addresses.

Exam trap

The trap here is assuming that allowing the VPC CIDR range is equivalent to allowing only the load balancer, when in fact it grants access to every resource in the VPC.

858
MCQhard

A healthcare analytics firm stores protected health information in an Amazon S3 bucket encrypted with SSE-KMS using a customer managed key. The firm's security team wants to ensure that only a specific IAM role used by an analytics application can decrypt objects, while other principals in the same account with broad S3 permissions cannot. The key policy currently grants kms:* to the account root. Which change should a solutions architect make to enforce the restriction?

A.Enable S3 Block Public Access on the bucket and require requests to use TLS by adding a bucket policy condition aws:SecureTransport.
B.Edit the KMS key policy to remove kms:Decrypt from the account root statement and add a statement that allows kms:Decrypt only to the analytics application's IAM role, then remove any identity-based policies that grant kms:Decrypt to other principals.
C.Add an S3 bucket policy that denies s3:GetObject to all principals except the analytics application's IAM role.
D.Change the bucket's default encryption from SSE-KMS to SSE-S3 so that AWS manages the keys and only the bucket owner can read objects.
AnswerB

KMS key policies are the primary access control for a customer managed key. Removing broad decrypt permissions from the root statement and granting kms:Decrypt only to the analytics role ensures that no other principal can decrypt, even if they have broad S3 permissions, because S3 decrypt operations require both S3 object access and kms:Decrypt on the key. Identity-based policies granting decrypt elsewhere must also be removed.

Why this answer

For SSE-KMS objects, reading an object requires permission to both the S3 object and the KMS key. The KMS key policy is the authoritative control that can limit kms:Decrypt to the analytics application's role. Removing broad decrypt grants from the root statement and from other identity-based policies ensures that other principals with S3 permissions cannot decrypt the protected health information.

Exam trap

The trap here is believing that an S3 bucket policy denying s3:GetObject fully protects SSE-KMS data, when the decisive control for decryption is the KMS key policy.

859
MCQhard

A company has a VPC with a CIDR block of 10.0.0.0/16. They need to deploy a web application that must be accessible from the internet. The application will run on Amazon EC2 instances in an Auto Scaling group. The security team requires that the instances be in private subnets and that inbound traffic from the internet be allowed only on ports 80 and 443. They also want to use an Application Load Balancer (ALB) for load balancing and SSL termination. Which architecture meets these requirements?

A.Deploy the ALB in public subnets, the EC2 instances in private subnets, and configure the ALB security group to allow inbound 80/443 from the VPC CIDR. Configure the EC2 security group to allow inbound traffic from the ALB security group.
B.Deploy the ALB in public subnets, the EC2 instances in public subnets, and configure both security groups to allow inbound 80/443 from 0.0.0.0/0. Configure the EC2 instances to use an Elastic IP address.
C.Deploy the ALB in private subnets, the EC2 instances in public subnets, and configure the ALB security group to allow inbound 80/443 from 0.0.0.0/0. Configure the EC2 security group to allow inbound traffic only from the ALB security group.
D.Deploy the ALB in public subnets, the EC2 instances in private subnets, and configure the ALB security group to allow inbound 80/443 from 0.0.0.0/0. Configure the EC2 security group to allow inbound traffic only from the ALB security group.
AnswerD

This architecture places the ALB in public subnets to receive internet traffic, while the EC2 instances remain in private subnets without public IPs. The ALB security group allows inbound 80/443 from the internet, and the EC2 security group allows inbound only from the ALB, ensuring that instances are not directly accessible. This meets all requirements and follows AWS best practices.

Why this answer

The correct architecture uses public subnets for the ALB and private subnets for the EC2 instances. The ALB security group allows inbound internet traffic on 80/443, while the EC2 security group allows inbound only from the ALB. This design ensures the instances are not directly accessible from the internet and meets the load balancing and SSL termination requirements.

Exam trap

The trap here is placing the ALB in private subnets or the instances in public subnets, which breaks the security requirements.

860
MCQmedium

A developer accidentally deletes important rows in an RDS database. The mistake is discovered 45 minutes later. The database has automated backups enabled with a retention period of 7 days. What is the best way to restore the database to a point just before the deletion?

A.Restore the latest manual snapshot and then run SQL scripts to revert the deletion.
B.Use point-in-time restore (PITR) to restore the database to a specific timestamp before the deletion, based on automated backups.
C.Promote an existing read replica to be the primary and then copy the missing rows from logs.
D.Recreate the instance using the most recent CloudWatch metric alarm snapshot of storage metrics.
AnswerB

With automated backups enabled, RDS supports PITR within the retention window. PITR lets you restore to any second within that window, so you can select a timestamp just before the destructive deletion occurred. This avoids restoring a potentially stale snapshot and eliminates the need for risky manual compensating scripts.

Why this answer

Point-in-time restore (PITR) allows you to restore an RDS DB instance to any second within the automated backup retention period (here, 7 days). Since the deletion occurred 45 minutes ago, you can specify a timestamp just before the deletion, and RDS will replay the transaction logs to bring the database to that exact state. This is the most precise and efficient recovery method for accidental data modifications.

Exam trap

The trap here is that candidates may assume manual snapshots or read replicas can be used for granular point-in-time recovery, but only automated backups with transaction logs enable restoring to a specific second within the retention period.

How to eliminate wrong answers

Option A is wrong because manual snapshots capture the entire instance at a point in time, but they do not provide the granularity to restore to a specific moment just before the deletion; you would lose all changes made after the snapshot, and running SQL scripts to revert deletions is error-prone and not a built-in RDS feature. Option C is wrong because promoting a read replica makes it a new primary, but it does not revert data; it simply becomes a writable copy of the current state, which still contains the deletion. Option D is wrong because CloudWatch metric alarms monitor performance metrics, not database row-level data; they cannot be used to restore or recover deleted rows.

861
MCQhard

Based on the exhibit, a partner account uploads encrypted objects to a central S3 bucket and later reads them back. The S3 permissions are correct, but the requests still fail. What change is required so the partner workload can use the customer-managed KMS key safely?

A.Replace SSE-KMS with S3 object ACLs so the partner account can bypass KMS authorization.
B.Create a new bucket in the partner account and copy the objects there to avoid cross-account encryption.
C.Switch the bucket to SSE-S3 so the partner role no longer needs KMS permissions.
D.Update the CMK key policy, or add a tightly scoped grant, to allow the partner role the required KMS actions through S3.
AnswerD

Cross-account access to SSE-KMS encrypted objects requires KMS authorization in addition to S3 authorization. The key policy must trust the partner role, and the permissions should be limited to the needed KMS actions such as Decrypt, Encrypt, and GenerateDataKey with a service condition for S3. That is why the partner can have valid S3 permissions and still fail until the KMS policy is fixed.

Why this answer

When using a customer-managed KMS key (CMK) for SSE-KMS in a cross-account scenario, the key policy must explicitly grant the partner account's IAM role the necessary KMS actions (kms:Decrypt, kms:GenerateDataKey) to allow S3 to perform the encryption/decryption on behalf of the partner. Without this policy update or a tightly scoped grant, S3 cannot authorize the KMS operation even if the S3 bucket policy permits the upload/read.

Exam trap

The trap here is that candidates assume the S3 bucket policy alone is sufficient for cross-account access with SSE-KMS, forgetting that KMS requires its own separate authorization via the key policy or a grant, which is a frequent point of failure in multi-account architectures.

Why the other options are wrong

A

S3 object ACLs do not bypass KMS authorization; the partner workload still needs KMS permissions to decrypt objects encrypted with a customer-managed KMS key. ACLs only control S3-level access, not encryption key access.

B

Creating a new bucket in the partner account and copying objects there does not solve the cross-account KMS authorization issue; the partner still needs KMS permissions to decrypt objects encrypted with the customer-managed KMS key, and moving objects does not grant those permissions.

C

Switching to SSE-S3 would remove KMS encryption, but the question states the objects are encrypted with a customer-managed KMS key (CMK). Changing encryption type is not a valid fix for KMS authorization issues; the partner workload must use the same CMK to read back the objects.

When would these options actually be correct?

A

This would be correct if the question asked for a way to grant cross-account access to S3 objects without encryption, or if the objects were not encrypted and the issue was solely about S3 permissions.

B

This option would be correct if the question required isolating data to avoid cross-account access entirely, such as when regulatory compliance mandates that data must not leave the partner's account, or when the central bucket policy cannot be modified to grant cross-account permissions.

C

This option would be correct if the question specified that the objects are encrypted with SSE-KMS (AWS managed key) and the partner workload does not need KMS permissions, or if the requirement is to simplify encryption and the objects are not sensitive. For example, a scenario where cost reduction or eliminating KMS key management is the goal.

Why candidates pick the wrong answer

A

Candidates may think ACLs provide a simpler alternative to KMS policies, misunderstanding that ACLs cannot override KMS encryption requirements.

B

Candidates may think that moving objects to the partner's own bucket eliminates the need for cross-account KMS permissions, overlooking that the objects are already encrypted with the customer's KMS key and decryption still requires access to that key.

C

Candidates may think switching to SSE-S3 is a quick way to bypass KMS permission issues without understanding that it changes the encryption method and may violate security requirements. They might also assume that S3 handles all encryption authorization automatically.

862
MCQhard

A claims portal uses Amazon RDS for PostgreSQL. Application credentials must not be stored on the EC2 instances, and authentication should use short-lived credentials. What should the architect recommend? The design must avoid adding custom operational scripts.

A.Store the database password in user data
B.Embed the database password in the AMI
C.IAM database authentication for RDS with an EC2 instance role
D.Use a security group rule that allows only application instances
AnswerC

IAM database authentication for RDS allows an EC2 instance to use its instance role to request a short-lived authentication token from AWS STS and present that token to PostgreSQL as the password. Because the credentials are temporary and scoped by IAM policies, there is no persistent database password stored in the application, and rotation is handled automatically when the token expires. This approach also lets you control access per IAM role or user, exactly the application's identity, while still requiring standard PostgreSQL authorization after authentication succeeds.

Why this answer

IAM database authentication for RDS PostgreSQL allows EC2 instances to authenticate using short-lived credentials (tokens) obtained via the IAM instance profile role, eliminating the need to store long-term credentials on the instance. The EC2 instance assumes an IAM role, which grants permission to generate an authentication token (valid for 15 minutes) using the AWS CLI or SDK, and that token is used as the password for the database connection. This approach satisfies the requirements of no stored credentials, short-lived authentication, and no custom operational scripts.

Exam trap

The trap here is that candidates often confuse network-layer controls (security groups) with application-layer authentication, or they assume that storing credentials in user data or an AMI is acceptable because it's 'not on the instance filesystem' — but both still persist the credential on the instance, violating the 'not stored on EC2' requirement.

How to eliminate wrong answers

Option A is wrong because storing the database password in user data means the password is written to the instance's metadata and remains on the instance, violating the requirement that credentials must not be stored on EC2 instances. Option B is wrong because embedding the database password in the AMI hard-codes the credential into the image, which persists across instance launches and again stores credentials on the instance, failing the no-storage requirement. Option D is wrong because a security group rule controls network access at the transport layer (IP/port) and does not provide authentication; it cannot replace database credentials or enforce short-lived authentication.

863
MCQmedium

A payments API uses an RDS MySQL database and must remain available during an Availability Zone failure with minimal application changes. What should the architect enable?

A.S3 Cross-Region Replication
B.Multi-AZ deployment for the RDS DB instance
C.Read replicas only
D.EBS snapshots every hour
AnswerB

A Multi-AZ deployment creates a synchronized standby replica in a different Availability Zone of the same Region. Amazon RDS automatically fails over to the standby when a problem is detected on the primary, providing high availability for the DB instance. This is the correct solution for automatic, synchronous failover with minimal data loss.

Why this answer

Multi-AZ deployment for RDS MySQL automatically provisions and maintains a synchronous standby replica in a different Availability Zone. In the event of an AZ failure, Amazon RDS automatically fails over to the standby, providing high availability with minimal application changes (the application only needs to reconnect to the same endpoint). This meets the requirement for availability during an AZ outage without requiring code modifications.

Exam trap

The trap here is that candidates often confuse read replicas (which are for read scaling and manual promotion) with Multi-AZ (which provides automatic failover for high availability), leading them to select read replicas as a cheaper or simpler alternative.

How to eliminate wrong answers

Option A is wrong because S3 Cross-Region Replication is for object storage in S3, not for RDS MySQL databases, and it does not provide automatic failover for a relational database. Option C is wrong because read replicas are designed for read scaling, not for automatic failover during an AZ failure; they require manual promotion and application changes to redirect writes. Option D is wrong because EBS snapshots every hour provide point-in-time backup and recovery, not high availability; restoring from a snapshot would involve significant downtime and manual intervention, not minimal application changes.

864
MCQmedium

You host a public API using Amazon API Gateway in two AWS Regions: us-east-1 (primary) and us-west-2 (secondary). You want Route 53 to send client traffic to the secondary region only when the primary API is unhealthy. Which Route 53 setup best meets this requirement?

A.Use latency-based routing with one routing policy per region, and use CloudWatch alarms to update traffic weights between regions.
B.Use Route 53 failover routing with two ALIAS records (same DNS name) pointing to the API Gateway regional endpoints: one record is configured as PRIMARY with an associated health check, and the other is configured as SECONDARY.
C.Use weighted routing across both regions and rely on Route 53 health checks to automatically set the secondary to 100% weight when the primary fails.
D.Use geolocation routing to map some client geographies to the secondary region and the rest to the primary region.
AnswerB

Route 53 failover routing is purpose-built for active-passive architecture across two endpoints. The PRIMARY ALIAS record is tied to a health check that monitors the API Gateway regional endpoint in the primary region; when that health check fails, Route 53 automatically returns the SECONDARY record's endpoint, redirecting traffic to the secondary region. This is the correct, fully managed way to achieve automatic regional failover for public APIs.

Why this answer

Route 53 failover routing is designed for active-passive setups where traffic is sent to a primary resource unless it is unhealthy, in which case traffic is routed to a secondary resource. By creating two ALIAS records with the same DNS name, one marked PRIMARY with an associated health check and the other marked SECONDARY, Route 53 will automatically fail over to the secondary region when the health check for the primary API Gateway endpoint fails. This directly meets the requirement of sending traffic to the secondary region only when the primary API is unhealthy.

Exam trap

The trap here is that candidates often confuse weighted routing with failover routing, mistakenly believing that Route 53 health checks can automatically adjust weights to achieve active-passive failover, when in fact weighted routing does not support dynamic weight adjustment based on health.

How to eliminate wrong answers

Option A is wrong because latency-based routing directs traffic based on lowest latency, not health, and using CloudWatch alarms to manually update weights is not an automatic failover mechanism; it also requires custom automation and does not natively support health-check-driven failover. Option C is wrong because weighted routing distributes traffic based on assigned weights and does not automatically set the secondary to 100% weight when the primary fails; Route 53 health checks can mark a record as unhealthy but do not dynamically adjust weights—they would cause the primary record to be excluded from responses, but the secondary would only receive traffic if its weight is non-zero, and the behavior is not a clean active-passive failover. Option D is wrong because geolocation routing directs traffic based on the geographic location of the client, not the health of the endpoint, and it cannot automatically fail over traffic from one region to another when the primary becomes unhealthy.

865
MCQmedium

A DynamoDB table uses this schema: partition key = customerId, sort key = timestamp. During a marketing campaign, one customer generates extremely high read traffic and the application sees ProvisionedThroughputExceeded errors even though the table’s total capacity is sufficient. What change most directly improves read distribution across partitions?

A.Increase the table’s provisioned read capacity units while keeping partition key = customerId.
B.Add a salt component to the partition key by changing it to customerId#salt, where salt is derived from a hash of requestId so a single customer’s requests are spread across many partitions; keep the sort key as timestamp.
C.Remove the sort key and use timestamp as the partition key to increase cardinality.
D.Switch to on-demand capacity and rely on DynamoDB to automatically distribute reads across partitions.
AnswerB

Hot partition throttling usually occurs when too many requests target a single partition key value. Salting transforms the partition key so that one high-traffic customerId maps to multiple distinct partition keys (e.g., customerId#0, customerId#1, etc.), which increases the number of partitions that can serve that customer’s workload concurrently and reduces the probability that a single partition becomes overloaded.

Why this answer

Adding a salt to the partition key (e.g., customerId#hash(requestId)) distributes the read-heavy customer's data across multiple physical partitions. This prevents a single hot partition from throttling requests, even when the table's total provisioned capacity is sufficient. DynamoDB's partition key determines the internal hash used for data placement, so increasing partition key cardinality directly improves read distribution.

Exam trap

The trap here is that candidates confuse total table capacity with per-partition capacity, assuming that increasing RCUs or switching to on-demand will fix throttling caused by a hot key, when in reality the bottleneck is the single partition's throughput limit.

Why the other options are wrong

A

Increasing read capacity units does not address the root cause: a single hot partition. The total capacity may be sufficient, but all reads for the hot customer hit the same partition, causing throttling at that partition level.

C

Changing the partition key to timestamp would cause all reads for a given time range to hit a single partition, creating a hot key and worsening the distribution issue, not solving it.

D

On-demand capacity handles throughput spikes but does not address the root cause: a single hot partition. The ProvisionedThroughputExceeded errors occur because one customer's data is concentrated on one partition, and on-demand capacity does not redistribute data across partitions.

When would these options actually be correct?

A

This option would be correct if the question described a scenario where overall table throughput is insufficient due to uniformly high traffic across all partitions, and the goal is simply to increase total capacity without a hot key issue.

C

If the question described a scenario where the access pattern is time-series with uniform read traffic across all timestamps (e.g., reading all records from the last hour for analytics), and the issue was low cardinality of the original partition key, then using timestamp as partition key could improve distribution.

D

This option would be correct in a scenario where the application experiences unpredictable, sudden traffic spikes across all partitions (e.g., viral social media campaign), and the issue is overall throughput capacity rather than a single hot key. On-demand capacity automatically scales to handle such spikes without manual provisioning.

Why candidates pick the wrong answer

A

Candidates often assume that insufficient capacity is the problem and that increasing RCUs will solve throttling, overlooking that DynamoDB throttles at the partition level, not the table level.

C

Candidates may think that increasing partition key cardinality always improves distribution, overlooking that timestamp as a partition key can create hot partitions for recent data.

D

Candidates may think on-demand capacity solves all throughput issues because it eliminates the need to manage capacity, overlooking that hot partitions are a data distribution problem that capacity alone cannot fix.

866
MCQeasy

A media company stores its video uploads in an Amazon S3 bucket. The security team wants to ensure that any objects uploaded to the bucket are encrypted at rest using keys managed in AWS Key Management Service (AWS KMS) and that the encryption key is rotated annually. Which solution should a solutions architect recommend?

A.Enable default encryption on the S3 bucket using SSE-KMS with a customer managed key configured for annual rotation.
B.Configure a bucket policy that denies uploads unless the request includes the x-amz-server-side-encryption header with the value AES256.
C.Enable default encryption on the S3 bucket using SSE-S3.
D.Use S3 server access logs to verify that objects are encrypted with a KMS key.
AnswerA

SSE-KMS with a customer managed key allows you to control the KMS key, and you can enable automatic key rotation with a period of one year. This meets the requirement for encryption at rest using KMS keys and annual rotation. S3 default encryption ensures all new objects are encrypted automatically.

Why this answer

Using SSE-KMS with a customer managed key satisfies both the encryption at rest requirement and the annual key rotation requirement. S3 default encryption automatically applies the chosen encryption method to all new objects, ensuring consistent protection without relying on upload headers. This approach provides the necessary control over the encryption keys.

Exam trap

The trap here is assuming that any encryption method satisfies a KMS-specific requirement, when only SSE-KMS with a customer managed key provides the required key management and rotation.

867
MCQmedium

A security requirement states: all uploads to an S3 bucket must (1) use TLS in transit and (2) use server-side encryption with AWS KMS (SSE-KMS) using the CMK key id 'abcd-1234'; otherwise the upload should be rejected. A developer reports that uploads are succeeding even though clients are sometimes using non-encrypted requests. Which bucket policy approach most directly enforces both controls?

A.Add an Allow statement granting s3:PutObject to the developer role; rely on IAM conditions in the developer role to enforce TLS and SSE-KMS.
B.Use Deny statements that reject PutObject when aws:SecureTransport is false and reject PutObject when s3:x-amz-server-side-encryption is not 'aws:kms' or when s3:x-amz-server-side-encryption-aws-kms-key-id does not equal 'abcd-1234'.
C.Enable S3 default encryption to SSE-KMS and remove any bucket policy enforcement, since default encryption automatically rejects all noncompliant uploads.
D.Attach a WAF rule to the S3 website endpoint to block non-TLS requests, because bucket policies cannot evaluate aws:SecureTransport.
AnswerB

These Deny conditions directly block noncompliant requests regardless of the caller’s IAM permissions because explicit Deny in a resource policy overrides any Allow. aws:SecureTransport identifies whether the request used TLS. The SSE-KMS headers (s3:x-amz-server-side-encryption and s3:x-amz-server-side-encryption-aws-kms-key-id) identify whether SSE-KMS was requested and which CMK key id was used.

Why this answer

Bucket policies can use the `aws:SecureTransport` condition key to enforce TLS and the `s3:x-amz-server-side-encryption` and `s3:x-amz-server-side-encryption-aws-kms-key-id` condition keys to enforce SSE-KMS with the specific CMK key ID. By using Deny statements, any request that does not meet both conditions is explicitly rejected, regardless of any Allow statements that might otherwise grant access. This directly enforces the security requirement at the bucket level.

Exam trap

The trap here is that candidates often confuse S3 default encryption with enforcement—default encryption only applies encryption to objects that lack it, but does not reject non-compliant uploads, so it cannot replace a bucket policy Deny statement for rejecting requests that violate encryption or TLS requirements.

How to eliminate wrong answers

Option A is wrong because relying on IAM conditions in the developer role does not enforce the controls for all clients; any client that can assume the role or use different credentials could bypass the conditions, and IAM conditions are not evaluated for anonymous or cross-account requests. Option C is wrong because S3 default encryption only applies server-side encryption to objects that are uploaded without an encryption header; it does not reject non-compliant uploads—it silently encrypts them, so requests without TLS or with a different KMS key ID would still succeed. Option D is wrong because AWS WAF cannot be attached directly to an S3 bucket endpoint; S3 does not support WAF integration, and bucket policies can indeed evaluate `aws:SecureTransport` to enforce TLS.

868
MCQhard

A media processing workflow in private subnets downloads large amounts of data from S3 through a NAT gateway. NAT data processing charges are high. What should the architect use to reduce cost?

A.S3 Object Lambda
B.AWS Shield Advanced
C.Gateway VPC endpoint for Amazon S3
D.A larger NAT gateway
AnswerC

A gateway VPC endpoint for Amazon S3 works by adding a prefix-list route to the VPC route table, directing S3-destined traffic from private subnets over the AWS backbone to the regional S3 endpoint. It requires no NAT gateway, no internet gateway, and no hourly charge, eliminating the per-GB NAT data processing fees on large media downloads. This is the correct cost-optimization mechanism for private-subnet access to S3.

Why this answer

A Gateway VPC endpoint for Amazon S3 allows instances in private subnets to access S3 directly via the AWS network without traversing a NAT gateway, eliminating NAT data processing charges. This is the most cost-effective solution because NAT gateway costs are incurred per GB of data processed, and using a gateway endpoint avoids those charges entirely.

Exam trap

The trap here is that candidates may think a larger NAT gateway would improve throughput and lower costs, but in reality, it only increases both hourly and per-GB charges, while a gateway VPC endpoint eliminates the data processing cost entirely.

How to eliminate wrong answers

Option A is wrong because S3 Object Lambda is used to transform data as it is retrieved from S3, not to reduce data transfer costs from private subnets. Option B is wrong because AWS Shield Advanced is a DDoS protection service that does not address NAT gateway data processing charges. Option D is wrong because a larger NAT gateway would increase, not reduce, costs due to higher hourly and data processing fees.

869
MCQmedium

An ECS service runs on EC2 instances and is fronted by an ALB. The ALB spans two Availability Zones, and the ECS service desired count is 2 tasks. The underlying EC2 capacity uses an Auto Scaling group (ASG) with min size set to 1, and the ASG also spans only one subnet in practice. What is the most effective change to meet the requirement that the service continues during a single-AZ instance loss?

A.Set the ECS deployment configuration to maximum percent 100 so tasks replace instances faster during rollouts.
B.Increase ASG min size to at least 2 and ensure the ASG uses subnets in at least two Availability Zones.
C.Enable ALB connection draining longer than expected so existing connections survive longer during an AZ event.
D.Reduce task memory reservations to pack both tasks onto a single EC2 instance.
AnswerB

Increasing the Auto Scaling group minimum to at least 2 and using subnets in at least two Availability Zones guarantees baseline EC2 capacity in multiple AZs. If one AZ becomes unavailable, the ALB can route traffic to healthy targets in the remaining AZ, and ECS has compute available to reschedule tasks. This directly provides the cross-AZ redundancy needed for the service to continue operating.

Why this answer

The current architecture has a single point of failure because the Auto Scaling group (ASG) spans only one subnet (one Availability Zone). If that AZ fails, all EC2 instances are lost, and the ECS service cannot run any tasks. Increasing the ASG min size to at least 2 and configuring it to use subnets in at least two AZs ensures that EC2 instances are distributed across AZs, allowing the ECS service to maintain at least one task in the surviving AZ during a single-AZ failure.

Exam trap

The trap here is that candidates often focus on ECS-specific settings (like deployment configuration or task placement) rather than recognizing that the root cause is the ASG's single-AZ limitation, which is a fundamental infrastructure resilience issue.

Why the other options are wrong

A

Setting maximum percent to 100 does not address the lack of multi-AZ redundancy; it only affects deployment speed, not availability during an AZ failure.

C

Connection draining helps preserve existing connections during a rolling update or instance deregistration, but it does not prevent service disruption when an entire Availability Zone fails. The ALB would still lose all healthy targets in that AZ, and new connections cannot be established to instances in the failed AZ.

D

Reducing task memory reservations does not address the single-AZ failure risk because both tasks could still be placed in the same Availability Zone, and the ASG only spans one subnet, so losing that AZ would still cause total service outage.

When would these options actually be correct?

A

If the question required faster task replacement during a rolling update to minimize downtime, setting maximum percent to 100 would allow new tasks to start before old ones are stopped, speeding up the deployment.

C

This option would be correct in a scenario where the question asks how to minimize disruption to in-flight requests during a planned deployment or when an instance is being replaced, and the ALB is configured to gradually drain connections before terminating instances.

D

This option would be correct in a scenario where the ECS service is running on Fargate with a single-AZ deployment and the goal is to reduce cost by packing tasks onto fewer instances, while high availability is not a requirement.

Why candidates pick the wrong answer

A

Candidates may confuse deployment configuration with high availability, thinking that faster task replacement compensates for missing AZ redundancy.

C

Candidates may think that extending connection draining provides enough time for the system to recover from an AZ failure, confusing it with a mechanism to handle abrupt instance loss rather than graceful termination.

D

Candidates may think that reducing resource reservations allows tasks to fit on fewer instances, which they mistakenly believe improves resilience by concentrating tasks, but they overlook the need for multi-AZ distribution to survive AZ failures.

870
MCQeasy

A web service runs on an Auto Scaling group (ASG). The team updates configuration (AMIs, environment variables) in a Launch Template and wants new instances created during scale-out to use the latest Launch Template version. What should the architect do?

A.Leave the ASG attached to the previous Launch Template version so scale-out is stable.
B.Set the ASG to use the latest Launch Template version and optionally start an instance refresh for existing instances.
C.Manually SSH into each new instance and reconfigure it after it launches.
D.Move the configuration changes into a security group rule so the ASG updates them automatically.
AnswerB

ASG scale-out uses the configured Launch Template version at instance launch time. Switching the ASG to the latest version ensures new instances are consistent. An instance refresh helps apply changes to running instances safely and predictably.

Why this answer

The ASG can be configured to use the latest version of a Launch Template by specifying the `$Latest` version alias. This ensures that any new instances launched during scale-out automatically use the most recent template configuration (e.g., updated AMI, environment variables). Additionally, an Instance Refresh can be triggered to roll the update across existing instances, aligning them with the same latest template version without manual intervention.

Exam trap

The trap here is that candidates may think the ASG automatically updates existing instances when the Launch Template version is changed, but without an Instance Refresh, only new scale-out instances receive the update, leaving existing instances on the old configuration.

Why the other options are wrong

C

Manually SSHing into each new instance is not scalable, error-prone, and violates automation best practices for Auto Scaling groups. The ASG should automatically use the latest Launch Template version to ensure new instances are correctly configured without manual intervention.

D

Security group rules control network traffic, not instance configuration like AMIs or environment variables. They cannot update ASG instances or launch templates.

When would these options actually be correct?

C

In a scenario where a legacy system requires one-time manual configuration changes on a few instances (e.g., applying a hotfix that cannot be automated via AMI or user data), and the ASG is not expected to scale frequently, manual SSH could be acceptable for immediate remediation.

D

If the question asked how to automatically allow new instances to receive traffic from a specific source after scaling out, updating a security group rule would be correct, as security groups apply to all instances in the ASG automatically.

Why candidates pick the wrong answer

C

Candidates may think manual intervention is a quick fix for configuration updates, especially if they are unfamiliar with Launch Template versioning and instance refresh features, or underestimate the operational overhead of manual steps at scale.

D

Candidates may confuse security groups with configuration management, thinking they can propagate changes to instances, or they may overestimate the scope of security group rules.

871
MCQmedium

A global mobile game backend serves mostly static images and JavaScript files from an S3 origin. Users in distant countries report slow load times. What should improve performance most?

A.RDS read replicas
B.Amazon CloudFront distribution with the S3 bucket as origin
C.A larger S3 bucket
D.An EC2 Auto Scaling group in one Region
AnswerB

A CloudFront distribution with the S3 bucket as its origin places static images at edge locations strategically positioned near end users, so each request is served from the nearest Point of Presence instead of traversing the entire Internet to the S3 origin. This dramatically lowers latency for global users, reduces the number of direct requests hitting S3, and can reduce origin data transfer costs. CloudFront also supports persistent connections and HTTP/2, which further improve the download experience for mobile clients.

Why this answer

Amazon CloudFront is a content delivery network (CDN) that caches static content (images, JavaScript files) at edge locations worldwide. By distributing content closer to users, it significantly reduces latency and improves load times for a global audience, making it the most effective solution for this use case.

Exam trap

The trap here is that candidates may confuse improving database read performance (RDS read replicas) with improving static content delivery, or assume that scaling compute resources (Auto Scaling) in a single Region can solve global latency issues.

How to eliminate wrong answers

Option A is wrong because RDS read replicas are designed to offload read traffic from a relational database, not to accelerate delivery of static files stored in S3. Option C is wrong because increasing the S3 bucket size does not affect network latency or data transfer speed; it only increases storage capacity. Option D is wrong because an EC2 Auto Scaling group in a single Region does not address global latency; it only provides scalability within that one Region, leaving distant users unaffected.

872
Multi-Selecthard

A fleet of test servers is rebuilt every week from AMIs. EBS volumes are often left behind after termination, and the team creates daily snapshots of every volume even when nothing changes. Which three actions most reduce storage cost while preserving recovery options? Select three.

Select 3 answers
A.Use gp3 for new EBS volumes instead of gp2 when similar performance is enough.
B.Automate snapshot creation and deletion with Amazon Data Lifecycle Manager.
C.Move old snapshots to the EBS Snapshot Archive tier once they are rarely restored.
D.Keep unattached volumes around for troubleshooting after instance termination.
E.Raise provisioned IOPS on every volume so snapshot restore time feels faster.
AnswersA, B, C

Switching from gp2 to gp3 decouples IOPS and throughput from volume capacity, so you pay only for the storage size needed rather than provisioning extra gigabytes to reach a performance tier. This directly cuts the per-GB cost of the test servers' volumes while preserving equivalent baseline performance and full snapshot recovery options.

Why this answer

Option A is correct because gp3 volumes are billed at a lower per-GB price than gp2 and let you provision IOPS and throughput independently of capacity, so workloads with similar performance needs cost less without sacrificing recovery capability. Option B is correct because Amazon Data Lifecycle Manager automates snapshot creation and, crucially, retention/deletion policies, which stops the accumulation of stale daily snapshots that drive storage cost. Option C is correct because the EBS Snapshot Archive tier stores rarely restored snapshots at a substantially lower price (up to ~75% cheaper) while still preserving them as a recovery option.

Option D is wrong because leaving unattached EBS volumes after termination incurs ongoing per-GB charges and directly increases cost rather than reducing it. Option E is wrong because raising provisioned IOPS increases cost and does not address the leftover volumes or redundant snapshots, and IOPS provisioning does not meaningfully speed snapshot restore.

Exam trap

The trap here is that candidates may think keeping unattached volumes is a valid recovery option, but it is more cost-effective to snapshot and delete them, and they may overlook that raising IOPS does not accelerate snapshot restore times.

873
Multi-Selectmedium

A solutions architect is designing a highly available relational database tier for a customer-facing order system that must survive the loss of an entire Availability Zone with minimal administrative effort and no application connection-string changes during failover. (Choose two.)

Select 2 answers
A.Create manual read replicas in each Availability Zone and repoint the application to a replica after failover
B.Enable a Multi-AZ DB cluster deployment for the RDS database so a writer and two readable standbys span three Availability Zones
C.Store the database on an EC2 instance with an EBS volume and take nightly snapshots to a second Availability Zone
D.Deploy the database as an Amazon RDS Multi-AZ DB instance so a standby is maintained in a second Availability Zone
E.Rely on the RDS automated backup retention window to restore the instance into a different Availability Zone during an outage
AnswersB, D

A Multi-AZ DB cluster runs one writer and two readable standbys across three Availability Zones and provides a single writer endpoint plus a reader endpoint. Failover is automatic and typically faster than a Multi-AZ DB instance, and the endpoint remains stable, so the application is unaffected. This meets the zone-failure and no-connection-string-change requirements.

Why this answer

Both Multi-AZ deployment types keep a synchronous standby copy of the data in another Availability Zone and expose a stable endpoint that survives automatic failover, which removes the need to edit connection strings. A Multi-AZ DB cluster goes further by adding two readable standbys across three zones and offering faster failover. The remaining choices rely on asynchronous replicas, manual promotion, or snapshot restores that all require operator action.

Exam trap

The trap here is assuming read replicas provide automatic failover; RDS read replicas replicate asynchronously and promoting one requires a manual step and can lose recent writes.

874
MCQmedium

A document portal requires consistent high IOPS for a transactional database on EC2. Which EBS volume type is most suitable?

A.sc1 Cold HDD
B.Instance store only
C.Provisioned IOPS SSD such as io2
D.st1 Throughput Optimized HDD
AnswerC

Provisioned IOPS SSD volumes such as io2 deliver consistently high, configurable IOPS independent of volume size, meeting the transactional database's steady performance demand. gp3 and gp2 burst and scale IOPS with capacity, so they cannot guarantee the sustained throughput required.

Why this answer

Provisioned IOPS SSD (io2) is the correct choice because it delivers consistent, high IOPS performance required for transactional databases running on EC2. io2 volumes offer a 99.999% durability and can sustain up to 256,000 IOPS per volume, making them ideal for latency-sensitive workloads like OLTP databases.

Exam trap

The trap here is that candidates often confuse 'high IOPS' with 'high throughput' and select st1 or sc1, not realizing that transactional databases require low-latency random I/O, which only SSD-based volumes like io2 can consistently deliver.

How to eliminate wrong answers

Option A is wrong because sc1 Cold HDD is designed for infrequently accessed, throughput-oriented workloads with low cost, and cannot provide consistent high IOPS due to its burst-bucket model and high latency. Option B is wrong because instance store volumes are ephemeral and data is lost on instance stop/termination, making them unsuitable for persistent transactional databases that require durability and consistent IOPS. Option D is wrong because st1 Throughput Optimized HDD is optimized for large, sequential workloads like big data and log processing, not for random I/O patterns typical of transactional databases, and its performance is limited to a maximum of 500 IOPS per volume.

875
MCQmedium

A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The design must avoid adding custom operational scripts.

A.Aurora Global Database
B.A single-AZ Aurora cluster
C.An ElastiCache Redis replica
D.Manual snapshots copied monthly
AnswerA

Aurora Global Database is the correct choice because it uses a primary Region plus up to five secondary Regions with dedicated storage-level replication typically under one second. This gives a low recovery point objective (RPO) and fast promoted read replica failover for disaster recovery, unlike snapshot-based or single-Region options. It also allows local reads in secondary Regions for low-latency access, directly meeting the ticket system's requirement for fast access and rapid DR.

Why this answer

Aurora Global Database is designed for cross-Region disaster recovery with a typical RPO of 1 second and RTO of 1 minute, using storage-based replication that does not require custom scripts. It replicates data from a primary Region to up to five secondary Regions with minimal impact on database performance, meeting the low RPO requirement without operational overhead.

Exam trap

The trap here is that candidates may confuse cross-Region read replicas (which require manual promotion and have higher RPO) with Aurora Global Database, which provides automated failover and lower RPO without custom scripts.

How to eliminate wrong answers

Option B is wrong because a single-AZ Aurora cluster lacks any cross-Region replication or failover capability, providing no disaster recovery across Regions. Option C is wrong because ElastiCache Redis is an in-memory cache, not a persistent database, and cannot serve as a primary data store for ticket bookings or provide cross-Region DR with low RPO. Option D is wrong because manual snapshots copied monthly result in an RPO of up to one month, which is far too high for fast disaster recovery, and the process requires custom scripting to automate cross-Region copy.

876
MCQeasy

A company runs a stateless containerized web application on Amazon ECS with the Fargate launch type behind an Application Load Balancer. The application must scale out quickly when request latency rises and scale in when traffic drops, and the operations team wants a managed target-tracking approach. Which solution meets these requirements?

A.Configure an Application Auto Scaling target-tracking scaling policy on the ECS service using the ALBRequestCountPerTarget metric.
B.Use a scheduled scaling action on the ECS service to add tasks at the start of each business day.
C.Create an EC2 Auto Scaling group with a target-tracking policy and register the ECS tasks as instances.
D.Deploy the application on EC2 instances in an Auto Scaling group and use the ECS-optimized AMI, replacing Fargate.
AnswerA

Application Auto Scaling supports target tracking for ECS services, and ALBRequestCountPerTarget is a predefined metric that scales tasks based on the load each task receives from the load balancer. It responds to demand changes automatically, works with Fargate because no EC2 capacity needs to be managed, and requires no custom scaling logic.

Why this answer

Application Auto Scaling target tracking on the ECS service is the managed mechanism that adjusts desired task count in response to a metric. ALBRequestCountPerTarget reflects how much work each task is receiving from the load balancer, so it scales out before latency degrades and scales in as traffic falls, and it works natively with the Fargate launch type.

Exam trap

The trap here is reaching for EC2 Auto Scaling group mechanics out of habit, when a Fargate service is scaled through Application Auto Scaling at the task level.

877
MCQhard

A healthcare data platform stores patient documents in an Amazon S3 bucket in us-east-1. Regulations require that the data remain readable with low latency even if the entire us-east-1 Region becomes unavailable, and that writes continue in a secondary Region. The team wants object-level replication with minimal operational overhead and must preserve version history. Which solution BEST meets these requirements?

A.Use S3 Transfer Acceleration on the existing bucket and rely on edge locations to serve reads during the outage.
B.Enable S3 Cross-Region Replication (CRR) with versioning enabled on both buckets, and configure the application to read from the replica bucket during a regional failure.
C.Enable S3 Same-Region Replication to a second bucket in us-east-1 and serve reads from that bucket during a regional failure.
D.Configure an S3 Lifecycle policy to transition objects to S3 Glacier Deep Archive in a second Region.
AnswerB

S3 CRR replicates objects asynchronously to a bucket in another Region and requires versioning on both source and destination, preserving version history. If us-east-1 becomes unavailable, the application can read from the replica bucket, and writes can be redirected there, satisfying the low-latency read and continued-write requirements.

Why this answer

Cross-Region Replication copies objects to a bucket in another Region and requires versioning on both buckets, which preserves version history. The replica bucket remains accessible if the primary Region fails, and writes can be directed there, meeting both the availability and data-durability goals with minimal operational effort.

Exam trap

The trap here is confusing S3 Transfer Acceleration or archival lifecycle rules with actual cross-Region data replication, which are unrelated features.

878
MCQhard

A mobile banking backend must ensure that only encrypted EBS volumes can be created in the account. What is the strongest preventive control?

A.Run a daily Lambda function to encrypt unencrypted volumes
B.Enable VPC Flow Logs
C.Tag encrypted volumes after creation
D.Use an SCP that denies ec2:CreateVolume when the encrypted condition is false
AnswerD

An SCP is an AWS Organizations policy that applies a permission guardrail to all IAM principals in member accounts, and it can deny ec2:CreateVolume when the request does not include the ec2:Encrypted condition key set to true. For example, a statement using the Bool condition on ec2:Encrypted: false will block any attempt to launch an unencrypted volume in every account under the organizational unit, regardless of IAM permissions. This is a preventive control that stops noncompliant resources before creation, which exactly satisfies the requirement, and an SCP cannot be overridden by user or role policies.

Why this answer

Service Control Policies (SCPs) are a preventive control that can deny the ec2:CreateVolume action when the encrypted condition is false. This ensures that only encrypted EBS volumes can be created, enforcing encryption at the point of request before any volume is provisioned. SCPs operate at the AWS Organizations level, making them the strongest preventive mechanism for account-wide enforcement.

Exam trap

The trap here is that candidates often confuse reactive controls (like Lambda remediation) with preventive controls (like SCPs), or they mistakenly think tagging or logging can enforce encryption requirements.

How to eliminate wrong answers

Option A is wrong because running a daily Lambda function to encrypt unencrypted volumes is a detective/reactive control, not preventive; it only remediates volumes after they have already been created unencrypted, violating the requirement to prevent creation in the first place. Option B is wrong because VPC Flow Logs capture network traffic metadata and have no ability to enforce or prevent EBS volume creation or encryption; they are a monitoring tool, not a preventive control. Option C is wrong because tagging encrypted volumes after creation is a labeling action that does not prevent unencrypted volumes from being created; it is a detective or organizational control, not a preventive one.

879
MCQmedium

Based on the exhibit, an administrator accidentally deleted data from Amazon RDS for PostgreSQL about 90 minutes ago. Which recovery approach best restores the database to the exact required point in time?

A.Restore the latest automated snapshot back onto the existing DB instance.
B.Restore the database to the specified point in time into a new DB instance.
C.Create a read replica and promote it after the deletion is noticed.
D.Enable Multi-AZ so the database can automatically undo application mistakes.
AnswerB

Point-in-time restore uses automated backups plus transaction logs to recreate the database at a specific moment. For accidental deletion, this is the correct RDS recovery method because it can recover the database to just before the bad change while preserving all legitimate data up to that point.

Why this answer

Amazon RDS for PostgreSQL supports Point-in-Time Recovery (PITR), which allows you to restore a DB instance to any second within the backup retention period, up to the last five minutes. Since the deletion occurred approximately 90 minutes ago, you can restore to that exact point in time by specifying the timestamp, and RDS will create a new DB instance from automated backups and transaction logs. This is the only option that recovers the exact state before the accidental deletion.

Exam trap

The trap here is that candidates confuse automated snapshots with point-in-time recovery, assuming a snapshot restore can target a specific time, when in fact snapshots are point-in-time captures and cannot replay transaction logs to reach an arbitrary second.

Why the other options are wrong

A

Restoring the latest automated snapshot would not recover data to a specific point in time 90 minutes ago; it only restores to the snapshot creation time, which is likely much earlier.

C

Creating a read replica does not restore data; it only creates a live copy of the current database, which already has the deletion. Promoting it would not recover the deleted data from 90 minutes ago.

D

Multi-AZ provides high availability through synchronous replication to a standby instance, but it does not enable point-in-time recovery or undo application mistakes like accidental data deletion.

When would these options actually be correct?

A

This option would be correct if the question asked to restore the database to the most recent available snapshot (e.g., after a catastrophic failure) and there was no need for point-in-time recovery.

C

This option would be correct if the question asked for a way to offload read traffic from the primary database or to provide a standby for disaster recovery in a different region, without needing to recover deleted data.

D

A question asks for the best approach to ensure automatic failover and minimize downtime during a planned maintenance event or an Availability Zone outage for an RDS database.

Why candidates pick the wrong answer

A

Candidates may think snapshots are the primary recovery mechanism and overlook that RDS supports point-in-time recovery for more granular restoration.

C

Candidates may think a read replica can be used for point-in-time recovery because it is a copy of the database, but they overlook that it reflects the current state, not a past state.

D

Candidates may mistakenly believe that Multi-AZ includes automatic backup or rollback capabilities, confusing high availability with data recovery features.

880
MCQmedium

A partner company needs read-only access to reports in an S3 bucket for a B2B file exchange site. The partner has its own AWS account. What is the most secure scalable access pattern? The design must avoid adding custom operational scripts.

A.Make the objects public and rely on difficult-to-guess object names
B.Create an IAM user in the company account and share the access keys
C.Create a bucket policy that grants the partner role least-privilege access to the required prefix
D.Copy the objects to a public website bucket
AnswerC

A bucket policy that sets the Principal element to the partner role's ARN (e.g., arn:aws:iam::123456789012:role/PartnerRole) and the Resource to backup/prefix/* grants read-only access to only that prefix while keeping the bucket private. This resource-based policy is evaluated together with the partner role's own identity-based policy, so both administrators can enforce least privilege, and the source account can revoke access centrally by editing the bucket policy. Using a role also leverages temporary credentials and CloudTrail auditing.

Why this answer

It uses a bucket policy with a principal ARN for the partner's AWS account, granting read-only access to a specific prefix. This is secure (no public exposure), scalable (no per-user credentials to manage), and avoids custom scripts by leveraging native AWS IAM and S3 policy evaluation. The partner can use their own IAM roles to access the bucket without sharing long-term access keys.

Exam trap

The trap here is that candidates may choose Option B (IAM user with shared keys) because it seems straightforward, but they overlook the security risk of long-term credentials and the operational burden of key rotation, which violates the 'most secure scalable' and 'avoid custom scripts' requirements.

How to eliminate wrong answers

Option A is wrong because making objects public relies on security through obscurity (difficult-to-guess names), which is not secure and violates the principle of least privilege; objects can be discovered via enumeration or leaks. Option B is wrong because creating an IAM user and sharing access keys introduces long-term credentials that must be securely rotated and managed, increasing operational overhead and risk of exposure, contradicting the 'avoid custom operational scripts' requirement. Option D is wrong because copying objects to a public website bucket makes them publicly accessible, losing all access control, and adds unnecessary data duplication and synchronization overhead.

881
MCQmedium

A global video platform serves mostly static images and JavaScript files from an S3 origin. Users in distant countries report slow load times. What should improve performance most?

A.A larger S3 bucket
B.Amazon CloudFront distribution with the S3 bucket as origin
C.RDS read replicas
D.An EC2 Auto Scaling group in one Region
AnswerB

CloudFront caches the static images and JavaScript at edge locations worldwide, so distant users fetch objects from nearby points of presence rather than the S3 origin, cutting latency. This directly addresses the slow load times reported in distant countries.

Why this answer

Amazon CloudFront is a content delivery network (CDN) that caches static content (images, JavaScript) at edge locations worldwide. By distributing content closer to users, it reduces latency and improves load times significantly compared to serving directly from a single S3 origin. This is the most effective solution for a global user base accessing static assets.

Exam trap

The trap here is that candidates might confuse 'scaling' (Auto Scaling, larger buckets) with 'latency reduction' (CDN), or mistakenly think database read replicas can serve static web assets, when in fact they are only for relational database read offloading.

How to eliminate wrong answers

Option A is wrong because S3 bucket size has no impact on performance; S3 scales automatically to handle any amount of data, and a larger bucket does not reduce latency for distant users. Option C is wrong because RDS read replicas are designed to offload read traffic from a relational database, not to serve static files like images or JavaScript; they address database query performance, not content delivery. Option D is wrong because an EC2 Auto Scaling group in one Region only scales compute capacity within that single geographic area, failing to reduce latency for users in distant countries who still must traverse long network paths.

882
MCQhard

A genomics research company runs a large-scale sequence alignment workload on AWS. The workload requires a shared file system that can be accessed concurrently by thousands of EC2 instances, provides high throughput and low latency, and supports POSIX permissions. The data set is about 500 TB and grows by 10 TB per month. The solutions architect needs to choose a storage solution that meets these performance and scalability requirements. Which solution should the architect use?

A.Amazon FSx for Lustre with a linked Amazon S3 data repository.
B.Amazon FSx for Windows File Server with Multi-AZ deployment.
C.Amazon EFS with Provisioned Throughput mode and a mount target in each Availability Zone.
D.Amazon S3 with S3 Transfer Acceleration and an S3 Gateway endpoint.
AnswerA

FSx for Lustre is a high-performance file system optimized for compute-intensive workloads like genomics. It provides sub-millisecond latencies, millions of IOPS, and hundreds of gigabytes per second of throughput. It supports POSIX permissions and can be linked to an S3 bucket for durable long-term storage. The file system can scale to hundreds of terabytes and is accessible from thousands of EC2 instances concurrently, making it ideal for this scenario.

Why this answer

Amazon FSx for Lustre is a fully managed, high-performance file system optimized for compute-intensive workloads such as genomics, machine learning, and HPC. It provides POSIX-compliant access, scales to hundreds of gigabytes per second throughput and millions of IOPS, and can be linked to an S3 bucket for durable storage. It supports thousands of concurrent EC2 instances, making it the best fit for large-scale sequence alignment.

Exam trap

The trap here is confusing general-purpose scalable file storage like Amazon EFS with a high-performance file system like FSx for Lustre, which is specifically designed for HPC and low-latency, high-throughput workloads.

883
MCQhard

A batch analytics job currently uses two NAT gateways in each of three Availability Zones, but only one private subnet per AZ needs outbound internet access. What should the architect review first? The design must avoid adding custom operational scripts.

A.Replacing every NAT gateway with an internet gateway attached to private subnets
B.Whether one NAT gateway per AZ is sufficient for the required private subnets
C.Disabling route tables
D.Moving all workloads to public subnets
AnswerB

NAT gateways are provisioned per Availability Zone, and the standard high-availability pattern is one NAT gateway per AZ, not multiple in the same AZ. If the batch analytics job runs in only one AZ or its required private subnets are all within a single AZ, a second NAT gateway in that same AZ provides no additional resilience while doubling the hourly and data-processing charges. The cost optimization is to confirm that each AZ contains at least one private subnet needing egress; if not, the NAT gateway count should be reduced to one per AZ with active private workloads.

Why this answer

The current setup uses two NAT gateways per AZ, which is likely over-provisioned and incurs unnecessary costs. Since only one private subnet per AZ requires outbound internet access, a single NAT gateway per AZ is typically sufficient to handle the traffic, and this is the first cost-optimization step to review before making other changes.

Exam trap

The trap here is that candidates may assume more NAT gateways are always better for high availability, but the question explicitly states only one private subnet per AZ needs outbound access, making a single NAT gateway per AZ the cost-optimized starting point.

How to eliminate wrong answers

Option A is wrong because internet gateways cannot be attached to private subnets; they are attached to VPCs and only work with public subnets that have a route to the IGW. Option C is wrong because disabling route tables would break all network connectivity, not just outbound internet access, and is not a valid optimization strategy. Option D is wrong because moving all workloads to public subnets would expose them directly to the internet, violating security best practices and the requirement to avoid custom operational scripts.

884
MCQeasy

You have EC2 instances in private subnets with no NAT gateway. They must retrieve secrets from AWS Secrets Manager without sending traffic to the public internet. Which VPC endpoint type is the correct choice for connecting to AWS Secrets Manager?

A.Create a Gateway VPC endpoint for Secrets Manager.
B.Create an Interface VPC endpoint (AWS PrivateLink) for Secrets Manager and associate security groups for the endpoint.
C.Use a Transit Gateway attachment to route traffic to the public internet for Secrets Manager.
D.Deploy a NAT gateway and allow outbound HTTPS traffic to Secrets Manager.
AnswerB

An interface VPC endpoint for Secrets Manager creates elastic network interfaces (ENIs) with private IPs inside your subnets, enabling direct, private access to the service over the AWS backbone. You can attach security groups to the endpoint to enforce inbound and outbound traffic controls, and endpoint policies further restrict which secrets or actions are allowed. This fully satisfies the requirement to avoid internet egress because no traffic leaves the AWS network.

Why this answer

AWS Secrets Manager is accessed via an API endpoint that uses HTTPS. Interface VPC endpoints (AWS PrivateLink) are the correct choice for connecting to services like Secrets Manager because they use elastic network interfaces (ENIs) with private IPs in your VPC, allowing traffic to stay within the AWS network. Gateway endpoints only support S3 and DynamoDB, not Secrets Manager.

Exam trap

The trap here is that candidates often confuse Gateway endpoints (which are free and only for S3/DynamoDB) with Interface endpoints (which incur hourly charges but support many services like Secrets Manager, KMS, and CloudWatch).

How to eliminate wrong answers

Option A is wrong because Gateway VPC endpoints only support Amazon S3 and DynamoDB, not AWS Secrets Manager. Option C is wrong because Transit Gateway attachments route traffic between VPCs and on-premises networks, but they do not provide private connectivity to AWS public services without a NAT gateway or internet gateway. Option D is wrong because deploying a NAT gateway would send traffic to the public internet, violating the requirement to avoid public internet traffic.

885
MCQeasy

A company wants to protect a critical application from a full Region outage. The secondary Region should keep only a small amount of infrastructure running most of the time to control cost. Which disaster recovery strategy fits best?

A.Pilot light
B.Active-active
C.Single-AZ deployment
D.Blue/green deployment
AnswerA

Pilot light is a disaster recovery pattern that keeps a minimal, low-cost core of the application running in the backup Region—often just the database, replication, and a small compute tier—so that in a full Region outage you can quickly scale out the remaining services by invoking pre-written IaC and orchestration runbooks. It deliberately avoids paying for idle production-scale capacity, trading a longer failover time (typically minutes to an hour) for significant cost savings while still preserving recovery capability.

Why this answer

The pilot light strategy is correct because it keeps a minimal core of infrastructure (e.g., a small database, a few EC2 instances) running in the secondary Region, while the bulk of the application remains dormant. In a full Region outage, the pilot light can be rapidly scaled up to full production capacity, meeting the requirement of low ongoing cost with the ability to recover from a complete Region failure.

Exam trap

The trap here is that candidates confuse 'pilot light' with 'active-active' or 'warm standby,' mistakenly thinking that any multi-Region setup must run full capacity, when the pilot light specifically minimizes cost by keeping only a minimal footprint until failover is triggered.

Why the other options are wrong

B

Active-active runs full production workloads in both regions simultaneously, which does not minimize infrastructure in the secondary region and increases cost, contrary to the requirement to keep only a small amount of infrastructure running most of the time.

C

Single-AZ deployment provides no protection against a full Region outage because it operates within a single Availability Zone. The question requires cross-Region disaster recovery, which this option does not address.

D

Blue/green deployment is a release strategy for updating applications with minimal downtime, not a disaster recovery strategy for region outages. It does not address infrastructure in a secondary region.

When would these options actually be correct?

B

Active-active would be correct for a scenario requiring zero downtime and immediate failover, where the application must serve traffic from multiple regions concurrently to handle high availability and low latency, and cost is not a primary constraint.

C

This option would be correct for a question asking for a cost-effective deployment that maximizes availability within a single Region, where the application can tolerate a single AZ failure and the goal is to minimize complexity and cost.

D

A company wants to deploy a new version of an application with zero downtime and the ability to quickly roll back if issues arise. Blue/green deployment would be the correct answer.

Why candidates pick the wrong answer

B

Candidates may choose active-active because it provides high availability and fast failover, but they overlook the cost implication of running full infrastructure in both regions, which contradicts the cost control requirement.

C

Candidates may think that deploying in a single AZ is sufficient for basic high availability, or they confuse AZ-level redundancy with Region-level disaster recovery, overlooking the requirement for cross-Region protection.

D

Candidates may confuse blue/green deployment with failover scenarios because both involve switching between environments, but blue/green is for updates, not disaster recovery.

886
MCQmedium

A static marketing site is served through CloudFront from an S3 origin. After a product update, customers report a drop in CloudFront cache hit ratio and the CloudFront bill increases because the origin is receiving many more requests for the same JS/CSS assets. Asset URLs are versioned, but requests now include an Authorization header even though these assets are public. Which CloudFront change most directly improves the cache hit ratio for these assets?

A.Increase the origin's max connections to handle more origin fetches
B.Configure the CloudFront cache policy so Authorization is not included in the cache key, and use an origin request policy that does not forward Authorization to the S3 origin for this behavior
C.Set CloudFront minimum TTL to 0 seconds so caches expire faster and origin fetches start again
D.Disable CloudFront compression because Authorization headers are not cacheable when compression is enabled
AnswerB

For public assets, Authorization should not vary the cache key. Removing it from the cache key allows CloudFront to reuse cached objects across requests, and not forwarding it to the origin avoids unnecessary origin variation and request overhead.

Why this answer

The drop in cache hit ratio is caused by the Authorization header being included in the cache key, which makes CloudFront treat each request as unique even when the asset URL is the same. By configuring the cache policy to exclude Authorization from the cache key and using an origin request policy that does not forward it to S3, CloudFront can serve cached responses for all users regardless of their Authorization header, restoring the cache hit ratio.

Exam trap

The trap here is that candidates may think increasing origin capacity or adjusting TTLs solves the problem, but the real issue is that the Authorization header is unnecessarily varying the cache key, which is a common misconfiguration in CloudFront when public assets are served alongside authenticated content.

How to eliminate wrong answers

Option A is wrong because increasing origin max connections addresses origin load but does not fix the root cause of cache misses caused by the Authorization header in the cache key. Option C is wrong because setting minimum TTL to 0 seconds forces CloudFront to revalidate every request with the origin, which would increase origin fetches and worsen the cache hit ratio and bill. Option D is wrong because CloudFront compression does not affect cacheability of Authorization headers; the header is simply not part of the cache key by default unless explicitly included, and disabling compression would not resolve the cache key issue.

887
MCQhard

A financial services company stores regulatory documents in an Amazon S3 bucket. The documents are accessed frequently for the first 90 days, then almost never, but must remain immediately retrievable for seven years. Retrieval latency of a few minutes is unacceptable, and the company wants the lowest storage cost that still meets the access requirement. Which S3 storage class should a solutions architect recommend?

A.S3 Glacier Flexible Retrieval
B.S3 One Zone-Infrequent Access (S3 One Zone-IA)
C.S3 Glacier Instant Retrieval
D.S3 Standard-Infrequent Access (S3 Standard-IA)
AnswerC

S3 Glacier Instant Retrieval is designed for long-lived archive data that is rarely accessed but requires millisecond retrieval. It costs less than S3 Standard-IA for long-term storage and meets the requirement that documents remain immediately retrievable for seven years. A lifecycle policy can transition objects after the active 90-day period.

Why this answer

The requirement is long-term, low-cost storage with immediate retrieval and no minutes-long latency. S3 Glacier Instant Retrieval provides millisecond access at a lower storage cost than S3 Standard-IA, making it the best fit for rarely accessed documents retained for years. A lifecycle rule can move objects from S3 Standard to Glacier Instant Retrieval after the active 90-day period.

Exam trap

The trap here is assuming any Glacier class introduces retrieval delay, when S3 Glacier Instant Retrieval actually provides millisecond access and is the intended low-cost class for archive data needing immediate retrieval.

888
MCQeasy

A company has a VPC with a CIDR block of 10.0.0.0/16. The company wants to allow its EC2 instances in a private subnet to access Amazon S3 without traversing the public internet. The company also wants to minimize data transfer costs. Which solution should a solutions architect recommend?

A.Configure an AWS Site-to-Site VPN connection to an on-premises data center that has a direct connection to Amazon S3.
B.Create a gateway VPC endpoint for Amazon S3 and update the route tables for the private subnets to use the endpoint.
C.Create an interface VPC endpoint for Amazon S3 in the VPC and update the route tables to use the endpoint.
D.Attach an internet gateway to the VPC and configure a NAT gateway in a public subnet to allow outbound traffic to Amazon S3.
AnswerB

Gateway VPC endpoints for Amazon S3 provide private connectivity without requiring an internet gateway or NAT device. They are free of charge and do not incur data processing fees. Updating the route tables to direct S3 traffic to the gateway endpoint ensures that traffic stays within the AWS network, meeting both the private access and cost minimization requirements.

Why this answer

A gateway VPC endpoint for Amazon S3 is the correct choice because it provides private connectivity from the VPC to S3 without using the public internet and is free of charge. Updating the route tables ensures that traffic to S3 is directed through the endpoint, meeting both the private access and cost minimization requirements.

Exam trap

The trap here is assuming that an interface VPC endpoint is required for private S3 access, when a gateway endpoint is the cost-effective and recommended solution for S3.

889
MCQhard

A healthcare company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The application must be accessible only to users connecting from a specific corporate IP range, and all traffic must be encrypted in transit. The security team wants to enforce these requirements at the load balancer level without modifying the application. Which combination of steps should a solutions architect take?

A.Configure an HTTPS listener on the ALB and use a network ACL on the public subnets to allow only the corporate IP range on port 443.
B.Configure an HTTPS listener on the ALB and enable AWS Shield Advanced to allow only the corporate IP range.
C.Configure an HTTP listener on the ALB and use AWS WAF to block requests from IPs outside the corporate range.
D.Configure an HTTPS listener on the ALB with an AWS Certificate Manager certificate, and attach a security group that allows inbound HTTPS only from the corporate IP range.
AnswerD

An HTTPS listener on the ALB encrypts traffic between clients and the load balancer using an ACM certificate. A security group attached to the ALB can restrict inbound access to the corporate CIDR range on port 443. This enforces both encryption in transit and source IP restriction without changing the application. The ALB terminates TLS, so backend traffic can be configured separately if needed.

Why this answer

To encrypt traffic in transit and restrict access to a corporate IP range at the ALB, use an HTTPS listener with an ACM certificate and a security group that allows inbound 443 only from the corporate CIDR. The ALB security group provides stateful, resource-level filtering. This meets both requirements without application changes and is the standard secure architecture for internet-facing ALBs.

Exam trap

The trap here is choosing AWS WAF or network ACLs for IP restriction while overlooking that an HTTP listener fails the encryption requirement.

890
Multi-Selectmedium

A startup runs an API on Amazon EC2. The instance must read items from one DynamoDB table and upload logs to one S3 bucket. Platform engineers also need a way to create new application roles, but those roles must never exceed a predefined set of permissions. Which three actions should the architect take? Select three.

Select 3 answers
A.Attach an IAM role to the EC2 instance profile and remove long-lived access keys from the server.
B.Give the EC2 instance an IAM user with administrator access for simplicity.
C.Scope the application policy to the exact DynamoDB table ARN and S3 bucket prefix.
D.Store the access keys in the application configuration file and rotate them later.
E.Use a permissions boundary for any IAM roles the platform team is allowed to create.
AnswersA, C, E

An IAM role attached to the EC2 instance profile supplies temporary credentials via instance metadata, so the application reads DynamoDB and writes to S3 without embedded long-lived access keys. Removing those keys eliminates a persistent credential that could be leaked or rotated manually.

Why this answer

Option A is correct because attaching an IAM role to the EC2 instance profile lets the instance obtain temporary credentials from the instance metadata service (IMDS) via AWS STS, eliminating the need for long-lived access keys on the server. Option C is correct because least-privilege policies should scope permissions to the specific resources needed — the exact DynamoDB table ARN and the S3 bucket/prefix ARN — rather than granting broad access. Option E is correct because a permissions boundary sets the maximum permissions an IAM role can have, so even if the platform team creates new application roles, those roles can never exceed the predefined boundary.

Option B is wrong because an IAM user with administrator access violates least privilege and embedding IAM user credentials on an instance is an anti-pattern. Option D is wrong because storing long-lived access keys in a configuration file is insecure and rotation does not remove the underlying risk that instance roles solve.

Exam trap

The trap here is that candidates may think storing access keys in a config file with rotation is acceptable, but AWS explicitly recommends using IAM roles for EC2 to avoid the security risks of long-lived static credentials.

891
MCQeasy

A public API is served through an Application Load Balancer and protected by AWS WAF. The team wants AWS to automatically block clients that send too many requests from the same IP address within a short time window. Which AWS WAF feature is the best fit?

A.Use a rate-based rule in AWS WAF to block when requests per IP exceed a configured threshold over the WAF rate-based evaluation window.
B.Use an AWS IAM policy on the ALB listener to deny requests when request count exceeds a threshold.
C.Enable S3 server access logs for the bucket that stores API responses and alert on high log volume.
D.Configure an AWS Lambda authorizer to reject requests after the Nth request from an IP address.
AnswerA

Rate-based rules are designed specifically to mitigate abusive traffic by limiting the number of requests from an identified source (typically by IP). When the threshold is exceeded, you can set the rule action to Block (or count first for tuning).

Why this answer

A rate-based rule in AWS WAF is specifically designed to automatically block clients when the number of requests from a single IP address exceeds a configured threshold within a rolling evaluation window (typically 5 minutes). This feature directly addresses the requirement to mitigate high request rates from the same IP, making it the best fit for the described use case.

Exam trap

The trap here is that candidates may confuse AWS WAF rate-based rules with other AWS services like IAM or Lambda authorizers, mistakenly thinking those can handle network-level rate limiting, when in fact only WAF provides native, automatic IP-based rate blocking.

How to eliminate wrong answers

Option B is wrong because IAM policies control authentication and authorization for AWS API calls, not network-level request rates on an ALB listener; they cannot deny HTTP requests based on request count. Option C is wrong because S3 server access logs are for auditing object-level access, not for real-time rate limiting of API requests, and alerting on log volume does not automatically block clients. Option D is wrong because an AWS Lambda authorizer is used for custom authentication/authorization of requests, not for rate limiting based on IP address; it would require custom logic and does not natively support sliding window rate tracking.

892
MCQmedium

A company hosts a web application on Amazon EC2 instances in a VPC. The application must access an Amazon RDS for MySQL database. The security team requires that database credentials never be stored in application code or on disk, and that credentials be automatically rotated every 30 days. Which solution should a solutions architect implement?

A.Store the credentials in an encrypted Amazon S3 object and have the application download and decrypt it at startup.
B.Store the credentials in AWS Secrets Manager and configure automatic rotation using the provided RDS rotation Lambda function.
C.Store the credentials in AWS Systems Manager Parameter Store as a SecureString parameter and enable automatic rotation with a custom Lambda function.
D.Use IAM database authentication for RDS and eliminate the need for a database password entirely.
AnswerB

Secrets Manager is designed for this use case: it stores secrets encrypted with KMS, and it offers built-in rotation for Amazon RDS through a managed Lambda function. The application retrieves credentials at runtime via the API or SDK, so nothing is stored in code or on disk, and rotation occurs automatically on the configured schedule.

Why this answer

AWS Secrets Manager provides managed storage and native rotation for RDS credentials. The application retrieves the secret at runtime, so no credentials are embedded in code or persisted on disk. The built-in rotation Lambda handles the 30-day schedule automatically, meeting both the security and operational requirements without custom code.

Exam trap

The trap here is assuming that Parameter Store SecureString parameters rotate automatically like Secrets Manager secrets, when Parameter Store requires custom rotation logic.

893
MCQmedium

A warehouse integration service receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers?

A.AWS WAF
B.Amazon Route 53 weighted routing
C.Amazon SQS queue
D.Amazon CloudFront
AnswerC

Amazon SQS is a fully managed message queue that decouples the warehouse integration service from order-producing applications. Producers enqueue each order, and consumers poll for messages at their own pace, allowing the queue to absorb sudden bursts and hold messages until processing capacity is available. Visibility timeout temporarily hides in-flight messages to prevent duplicate processing, while a dead-letter queue captures messages that repeatedly fail, enabling safe retries and preserving the order data during transient outages.

Why this answer

Amazon SQS is the correct choice because it acts as a durable, fully managed message queue that decouples the web tier from the fulfilment workers. It can absorb bursts of orders by storing messages durably, and workers can poll the queue at their own pace, with built-in retry logic via visibility timeouts and dead-letter queues to ensure no requests are lost.

Exam trap

The trap here is that candidates may confuse load-balancing or caching services (like Route 53 or CloudFront) with message queuing, failing to recognize that only a durable queue like SQS provides the necessary buffering, decoupling, and retry semantics for asynchronous order processing.

How to eliminate wrong answers

Option A is wrong because AWS WAF is a web application firewall that filters HTTP/S traffic based on rules, not a queuing or buffering mechanism; it cannot absorb spikes or retry processing. Option B is wrong because Amazon Route 53 weighted routing distributes DNS traffic across multiple endpoints based on weights, but it does not provide durable storage or retry capabilities for individual requests. Option D is wrong because Amazon CloudFront is a content delivery network (CDN) that caches static and dynamic content at edge locations; it can reduce load on origins but cannot queue or retry individual order messages.

894
MCQeasy

A workload runs in private subnets. It must access AWS services such as Amazon S3, but the company wants to avoid using a NAT Gateway to reduce outbound networking costs. What is the best solution?

A.Create VPC endpoints for the required AWS services and route traffic to them
B.Attach Elastic IP addresses to instances in private subnets
C.Install a NAT Gateway in every subnet to minimize routing hops
D.Open outbound internet access with a security group rule to reach service endpoints directly
AnswerA

VPC endpoints provide private connectivity from your VPC to supported AWS services without traversing the public internet or a NAT Gateway. For example, you can use a gateway endpoint for S3 (and interface endpoints for other services where supported), which avoids NAT Gateway hourly and data-processing charges.

Why this answer

VPC endpoints (Gateway Endpoints for S3 and DynamoDB, or Interface Endpoints for other services) allow instances in private subnets to access AWS services privately without traversing the internet or a NAT Gateway. This eliminates NAT Gateway data processing and hourly charges, directly reducing outbound networking costs while keeping traffic within the AWS network.

Exam trap

The trap here is that candidates often assume private subnets must use a NAT Gateway or internet gateway for any AWS service access, overlooking that VPC endpoints provide direct, cost-free connectivity to supported services within the AWS network.

Why the other options are wrong

B

Attaching Elastic IP addresses to instances in private subnets does not provide internet access because private subnets lack a route to an internet gateway; Elastic IPs require an internet gateway to be reachable.

C

Installing a NAT Gateway in every subnet increases costs (each NAT Gateway incurs hourly and data processing charges) and does not reduce outbound networking costs as required by the question.

D

Security group rules control inbound and outbound traffic at the instance level, but they cannot provide direct private connectivity to AWS services like S3. Instances in private subnets without a NAT Gateway or VPC Endpoint cannot reach public service endpoints over the internet.

When would these options actually be correct?

B

If the question required instances in a public subnet to have static public IP addresses for outbound internet access, attaching Elastic IPs would be correct, assuming an internet gateway is configured.

C

A question where the requirement is to provide outbound internet access to instances in private subnets for general internet access (not just AWS services), and cost is not a primary concern. For example: 'A workload in private subnets needs to download patches from the internet. What is the most reliable solution?'

D

If the question asked about allowing outbound internet access from instances in a public subnet to a specific IP range, and the goal was to restrict traffic at the instance level, then a security group rule would be appropriate. For example: 'A web server in a public subnet needs to access an external API; which security group configuration allows this?'

Why candidates pick the wrong answer

B

Candidates may mistakenly think Elastic IPs directly enable internet access, overlooking that private subnets cannot route to the internet without a NAT gateway or internet gateway.

C

Candidates may think that placing a NAT Gateway in each subnet reduces latency or routing hops, and they might overlook the cost implications, especially if they are focused on network performance rather than cost optimization.

D

Candidates may think that security groups can replace network infrastructure components like NAT Gateways or VPC Endpoints, misunderstanding that security groups only filter traffic but do not provide routing or private connectivity.

895
MCQmedium

A healthcare document service stores audit logs in S3. The compliance team requires that logs cannot be overwritten or deleted for seven years. What should be configured?

A.S3 Object Lock in compliance mode with an appropriate retention period
B.S3 server access logging
C.S3 lifecycle expiration after seven years
D.S3 versioning only
AnswerA

S3 Object Lock in compliance mode ensures that objects, once written, cannot be overwritten or deleted by any user, including the root account, until the specified retention period expires. This "write-once-read-many" (WORM) model is crucial for audit logs, providing an immutable record that meets stringent regulatory compliance requirements for data integrity and non-repudiation. It directly prevents any accidental or malicious alteration of the log data, making it tamper-proof.

Why this answer

S3 Object Lock in compliance mode prevents any user, including the root user, from overwriting or deleting objects for the specified retention period. This meets the compliance requirement of immutable audit logs for seven years, as compliance mode enforces a strict write-once-read-many (WORM) model that cannot be bypassed.

Exam trap

The trap here is that candidates often confuse S3 versioning with immutability, assuming versioning alone prevents deletion, but versioning only protects against accidental overwrites by creating new versions—it does not prevent explicit deletion of the current version or the entire object.

How to eliminate wrong answers

Option B is wrong because S3 server access logging only records requests made to the bucket, it does not prevent deletion or overwriting of existing logs. Option C is wrong because S3 lifecycle expiration automatically deletes objects after seven years, which violates the requirement that logs cannot be deleted. Option D is wrong because S3 versioning alone preserves previous versions but does not prevent deletion of the current version or overwriting of objects; it must be combined with Object Lock to enforce immutability.

896
MCQhard

An EC2 instance in a private subnet must access an S3 bucket that contains regulated exports for a image sharing application. The security team requires access to be allowed only when traffic comes through a specific VPC endpoint. What should the architect add to the bucket policy?

A.A condition that matches aws:sourceVpce to the endpoint ID
B.A deny statement for all IAM users except the EC2 role
C.A security group rule that allows HTTPS to S3
D.A condition that matches aws:RequestedRegion to the bucket Region
AnswerA

The aws:sourceVpce condition in a bucket policy restricts access to requests that arrive through a specific VPC endpoint, matching the endpoint ID exactly. This enforces the network path by verifying that the request source is the VPC endpoint you created for the private subnet, preventing access from the public internet or other VPCs. It is a robust, network-based condition that works alongside IAM identity policies, and it directly proves that traffic transited the endpoint.

Why this answer

The `aws:sourceVpce` condition key in an S3 bucket policy allows you to restrict access so that only traffic originating from a specific VPC endpoint (VPCe) is permitted. This enforces the security team's requirement that all S3 access must come through that endpoint, ensuring that requests from other paths (e.g., NAT gateway, internet gateway) are denied. The condition is evaluated at the S3 service side, not at the instance level, making it a direct and secure way to enforce the policy.

Exam trap

The trap here is that candidates often confuse `aws:sourceVpce` with `aws:SourceVpc` (which matches the VPC ID, not the endpoint ID) or incorrectly think a security group rule can enforce endpoint-specific routing, but security groups cannot control the network path taken by traffic.

How to eliminate wrong answers

Option B is wrong because denying all IAM users except the EC2 role does not enforce the requirement that traffic must come through a specific VPC endpoint; it only restricts which IAM identity can access the bucket, not the network path. Option C is wrong because a security group rule controls traffic at the instance level (allowing HTTPS to S3 from the instance), but it cannot enforce that the traffic must traverse a specific VPC endpoint—security groups do not have awareness of VPC endpoints. Option D is wrong because `aws:RequestedRegion` restricts the AWS Region from which the request is made, not the network path or VPC endpoint; it does not ensure traffic flows through the required VPC endpoint.

897
MCQmedium

A telemetry pipeline uses an Application Load Balancer in one Region. Global users need lower network latency to the application without caching dynamic responses. What should be considered? The architecture review board prefers a managed AWS-native control.

A.AWS Global Accelerator
B.S3 Cross-Region Replication
C.CloudFront only with long TTLs
D.AWS Backup cross-Region copy
AnswerA

Global Accelerator assigns two static anycast IP addresses and accepts TCP/UDP traffic at AWS edge locations, then forwards it across the AWS private global backbone to the configured Application Load Balancer. This avoids the public internet's variable routing hops and packet loss, which materially reduces latency and jitter for dynamic, non-cacheable telemetry traffic. It also supports weighted endpoint groups and automatic failover, making it the right performance layer for an ALB-backed pipeline.

Why this answer

AWS Global Accelerator is the correct choice because it uses the AWS global network and Anycast IP addresses to route user traffic to the optimal Application Load Balancer endpoint, reducing latency for global users without caching dynamic responses. Unlike CloudFront, Global Accelerator does not cache content; it simply optimizes the network path, making it ideal for dynamic or real-time applications where caching is not acceptable. It is a managed AWS-native service that aligns with the architecture review board's preference.

Exam trap

The trap here is that candidates often confuse CloudFront with Global Accelerator, assuming that any CDN-like service is the answer for latency reduction, but CloudFront's caching behavior makes it unsuitable for dynamic content that must not be cached.

How to eliminate wrong answers

Option B (S3 Cross-Region Replication) is wrong because it is designed for replicating objects in S3 buckets across regions for data durability or compliance, not for reducing network latency to an ALB-based application. Option C (CloudFront only with long TTLs) is wrong because CloudFront caches content at edge locations, which would cache dynamic responses—contradicting the requirement to avoid caching—and long TTLs would further exacerbate stale data issues. Option D (AWS Backup cross-Region copy) is wrong because it is a backup and disaster recovery service for creating copies of resources across regions, not a solution for improving application latency.

898
MCQeasy

A small e-commerce company hosts its product catalog on a single Amazon EC2 instance in a public subnet. Traffic is steady and predictable, and the instance runs 24/7. The company wants to reduce its monthly compute bill without changing the architecture or risking availability. Which action should a solutions architect recommend?

A.Move the instance to a Spot Instance to take advantage of lower hourly rates.
B.Enable detailed monitoring and create a CloudWatch alarm to stop the instance when traffic is low.
C.Purchase a one-year Standard Reserved Instance for the instance.
D.Migrate the workload to a smaller instance type and enable burstable performance.
AnswerC

A Standard Reserved Instance is ideal for a steady, always-on instance that will not change family or Region. It provides a significant discount over On-Demand in exchange for a one- or three-year commitment, and it requires no architectural change. This directly lowers the monthly compute bill while preserving availability.

Why this answer

For a steady, always-on instance that will not change family or Region, a Standard Reserved Instance delivers a substantial discount over On-Demand with no architectural change and no availability risk. It is the simplest, lowest-risk way to cut the monthly compute bill for this predictable workload.

Exam trap

The trap here is reaching for Spot because the hourly rate is lower, when the availability requirement rules out any interruptible purchasing option.

899
MCQmedium

A Lambda function for a order processing API needs to read a database password. The password must rotate automatically every 30 days and should not be stored in environment variables. Which service should be used?

A.AWS Secrets Manager with rotation enabled
B.An encrypted object in Amazon S3
C.AWS Systems Manager Parameter Store SecureString without automation
D.A KMS-encrypted Lambda environment variable
AnswerA

AWS Secrets Manager is the only option that provides a purpose-built, managed secret lifecycle with automatic rotation. Rotation is implemented by a custom or AWS-provided Lambda function that updates the secret value in the store and in the target database or service, ensuring credentials are periodically changed without application downtime. The service also supports fine-grained access policies, versioning during rotation, and native integration with Amazon RDS, Redshift, and other AWS services, so it is the optimal choice for an order-processing API that must rotate credentials automatically.

Why this answer

AWS Secrets Manager is designed to securely store, retrieve, and automatically rotate database credentials on a schedule. It natively supports rotation every 30 days via a built-in Lambda rotation function, and it avoids storing the password in environment variables, meeting both security and compliance requirements.

Exam trap

The trap here is that candidates often confuse AWS Systems Manager Parameter Store SecureString (which can store encrypted values but lacks automatic rotation) with Secrets Manager, or they assume KMS-encrypted environment variables are sufficient despite their static nature and the explicit requirement to avoid environment variables.

How to eliminate wrong answers

Option B is wrong because storing an encrypted object in Amazon S3 requires manual retrieval and decryption logic in the Lambda function, and it does not provide automated rotation every 30 days. Option C is wrong because AWS Systems Manager Parameter Store SecureString without automation lacks built-in rotation capabilities; you would need to implement custom rotation logic, which is not automatic. Option D is wrong because a KMS-encrypted Lambda environment variable is static and cannot be rotated automatically; any rotation would require redeploying the function, and the password remains in the environment variable, which is explicitly prohibited.

900
MCQhard

A financial analytics platform runs a stateless API on Amazon EC2 instances in an Auto Scaling group behind a Network Load Balancer. The API reads from an Amazon Aurora MySQL cluster that has a single writer instance and one reader instance in a different Availability Zone. During a recent Availability Zone event, the writer instance failed and the API saw several minutes of failed writes. The team wants writes to resume automatically with minimal downtime and no application code changes. What should the solutions architect do?

A.Ensure the Aurora cluster has at least one Aurora Replica in a different Availability Zone and rely on the cluster's automatic failover, which updates the writer endpoint to point at the promoted replica.
B.Add a second Aurora reader instance in a third Availability Zone and update the application connection string to use the reader endpoint for all queries.
C.Create an Aurora Replica in a second Availability Zone and promote it manually to writer using the AWS Management Console when a failure occurs.
D.Convert the cluster to an Aurora multi-master cluster so every instance can accept writes, and keep the existing writer endpoint in the connection string.
AnswerA

Aurora automatically promotes an Aurora Replica to writer when the current writer fails, and the cluster writer endpoint is repointed to the new writer, so the application keeps using the same endpoint and resumes writes without code changes. Having the replica in a separate Availability Zone ensures the failover target survives the zone event.

Why this answer

Aurora provides automatic failover by promoting an existing Aurora Replica to the writer role and updating the cluster writer endpoint, so applications that connect through that endpoint reconnect without modification. The prerequisite is that a replica exists in a different Availability Zone from the writer, which makes it a viable failover target when the writer's zone is lost.

Exam trap

The trap here is reaching for manual promotion or multi-master when Aurora already performs automatic failover to a replica, provided the replica exists in a different Availability Zone.

Page 11

Page 12 of 13

Page 13