Courseiva

SAA-C03 (SAA-C03) — Questions 376–450

935 questions total · 13pages · All types, answers revealed

Page 5

Page 6 of 13

Page 7
376
MCQmedium

A payments platform requires disaster recovery across Regions. Requirements: RPO of 15 minutes and RTO of about 1 hour. The business cannot afford full duplicate capacity in both Regions all the time, but the team wants automated readiness so failover is mostly operationally guided rather than a slow rebuild. Which DR strategy is the best fit?

A.Backup and restore only, relying on scheduled snapshots and manual restores during incidents.
B.Pilot light, keeping only minimal infrastructure in the secondary Region and starting full services after failover.
C.Warm standby, keeping core infrastructure and a partially provisioned environment ready in the secondary Region with frequent data replication.
D.Active/active, routing production traffic to both Regions continuously and accepting dual-region complexity.
AnswerC

Warm standby keeps core services and a scaled-down but functional environment running in the secondary Region with continuous replication, achieving the 15-minute RPO and roughly one-hour RTO. It avoids the cost of full duplicate capacity while remaining far faster than a rebuild.

Why this answer

Warm standby is the best fit because it maintains a partially provisioned environment in the secondary Region with core infrastructure (e.g., a smaller EC2 instance fleet, a replicated database) and uses frequent data replication (e.g., Amazon RDS cross-Region replication or DynamoDB global tables) to achieve an RPO of 15 minutes. The RTO of about 1 hour is achievable by scaling up the standby environment and redirecting traffic, which is faster than a full rebuild but avoids the cost of full duplicate capacity. This balances the business constraint of not affording active/active with the need for automated readiness and guided failover.

Exam trap

The trap here is that candidates often confuse pilot light with warm standby, assuming minimal infrastructure is sufficient for a 1-hour RTO, but pilot light requires provisioning compute resources after failover, which adds significant time, whereas warm standby already has compute running and only needs scaling.

How to eliminate wrong answers

Option A is wrong because backup and restore only relies on scheduled snapshots (e.g., EBS snapshots or RDS automated backups) and manual restores, which typically cannot achieve an RPO of 15 minutes (snapshots are often taken every few hours) and would result in an RTO far exceeding 1 hour due to manual intervention and data restoration time. Option B is wrong because pilot light keeps only minimal infrastructure (e.g., a small database replica and no application servers) in the secondary Region, and starting full services after failover requires provisioning compute resources, which would likely exceed the 1-hour RTO target. Option D is wrong because active/active requires full duplicate capacity in both Regions all the time, which contradicts the business constraint that they cannot afford this, and it introduces dual-region complexity that is unnecessary for the stated RPO/RTO goals.

377
MCQeasy

Your application uses ElastiCache Redis as a cache for user profiles stored in DynamoDB. You must ensure that when a profile is updated, subsequent reads see the latest value quickly. Which cache strategy is generally the best fit for this requirement?

A.Write to DynamoDB only, and never update or invalidate the Redis cache.
B.Use a cache-aside approach with TTL plus explicit invalidation after writes.
C.Cache only for reads, and do not fetch from DynamoDB when a key is missing.
D.Rely on eventual consistency of Redis replication to propagate updates to all nodes.
AnswerB

A cache-aside (lazy loading) pattern reads from cache first; if missing/expired, it fetches from the source of truth. After an update, explicitly invalidating or updating the cached entry ensures subsequent reads quickly reflect changes. TTL provides protection against missed invalidations while invalidation accelerates correctness after writes.

Why this answer

A cache-aside (lazy loading) strategy with TTL and explicit invalidation ensures that after a write to DynamoDB, the stale Redis entry is removed, forcing the next read to fetch the fresh profile from DynamoDB and repopulate the cache. This combination minimizes the window of stale reads while maintaining high read performance, which is critical for user profile caches where consistency matters.

Exam trap

The trap here is that candidates often confuse eventual consistency within Redis replication (which only applies to Redis-to-Redis sync) with the need to synchronize the cache with the authoritative data store (DynamoDB), leading them to pick option D, which does not address the core requirement of reflecting DynamoDB updates in the cache.

How to eliminate wrong answers

Option A is wrong because never updating or invalidating the Redis cache means stale data persists indefinitely, violating the requirement that subsequent reads see the latest value quickly. Option C is wrong because caching only for reads and not fetching from DynamoDB when a key is missing would result in cache misses returning no data, effectively breaking the application's ability to serve user profiles. Option D is wrong because relying on eventual consistency of Redis replication does not guarantee that updates to DynamoDB are reflected in the cache; Redis replication only synchronizes data between Redis nodes, not between DynamoDB and Redis, and does not address cache invalidation after writes.

378
MCQhard

A financial services firm runs a latency-sensitive trading application on Amazon EC2 instances distributed across three Availability Zones behind a Network Load Balancer. The application must continue serving traffic with no manual intervention if an entire Availability Zone becomes impaired, and each instance must receive a fair share of connections. Which combination of features meets these requirements?

A.Enable cross-zone load balancing on the Network Load Balancer and register targets in all three Availability Zones
B.Use an Application Load Balancer with sticky sessions enabled and register targets in all three Availability Zones
C.Attach an Elastic IP address to each instance and have clients connect directly using a published list of addresses
D.Create a Network Load Balancer with one target group per Availability Zone and rely on DNS failover between the groups
AnswerA

Cross-zone load balancing on a Network Load Balancer distributes traffic evenly across registered targets in every enabled Availability Zone, so a single impaired zone does not concentrate load on the surviving instances. With targets registered in all three zones, the load balancer's health checks automatically remove unhealthy targets and continue routing to the rest. This delivers both even connection distribution and unattended zone-failure tolerance.

Why this answer

Cross-zone load balancing on a Network Load Balancer spreads incoming connections evenly across registered targets in all enabled Availability Zones, which is exactly what the fair-share requirement demands. Registering targets in three zones lets health checks detect an impaired zone and remove its targets so the remaining instances absorb traffic without operator action. The other options either bypass health checking, bind clients to single targets, or rely on slow DNS-based failover.

Exam trap

The trap here is assuming that a Network Load Balancer distributes traffic evenly across zones by default; cross-zone load balancing must be enabled, and forgetting it can leave one zone carrying a disproportionate share.

379
MCQhard

A company runs a stateful web application on a fleet of Amazon EC2 instances in an Auto Scaling group. The application stores session state locally on each instance. During an Availability Zone failure, the Auto Scaling group replaces the unhealthy instances in a different AZ, but users lose their sessions and must log in again. The company wants to make the application resilient to AZ failures without requiring users to re-authenticate. Which solution should a solutions architect recommend?

A.Use Amazon S3 to store session data and configure the application to read and write session files directly to an S3 bucket.
B.Configure the Auto Scaling group to launch instances in only one Availability Zone and use a larger instance type to handle the load.
C.Store session state in an Amazon ElastiCache for Redis cluster with Multi-AZ enabled, and configure the application to use it.
D.Enable sticky sessions on the Application Load Balancer and increase the Auto Scaling group's desired capacity.
AnswerC

ElastiCache for Redis with Multi-AZ provides automatic failover to a replica in another AZ, and the session data is externalized from the EC2 instances. When an instance is replaced, it can retrieve the session from the Redis cluster, so users remain authenticated. This design decouples session state from the compute layer and survives an AZ failure.

Why this answer

Externalizing session state to a Multi-AZ ElastiCache for Redis cluster removes the dependency on local instance storage. When an instance is replaced after an AZ failure, the new instance can retrieve the session from Redis, so users stay logged in. This decouples state from compute and provides automatic failover, meeting the resilience requirement.

Exam trap

The trap here is thinking that sticky sessions or increased capacity can preserve session state when an Availability Zone fails.

380
MCQeasy

An ECS service runs on EC2 capacity. During peak traffic, tasks frequently wait for available container instances. The team wants faster scale-out for the underlying EC2 capacity when tasks increase. What is the best first architectural step?

A.Tune the container health check settings so tasks stop failing and stay running.
B.Use an ECS capacity provider (or Auto Scaling integration) to scale the EC2 instances based on ECS demand.
C.Pin all tasks to a single Availability Zone to reduce placement overhead.
D.Switch the tasks to run only on Fargate so EC2 scaling is no longer relevant.
AnswerB

When ECS tasks need compute, capacity must scale at the EC2 layer so there are enough container instances to place tasks. Integrating ECS with an Auto Scaling capacity provider allows the cluster to scale out in response to pending tasks. This reduces waiting time and improves responsiveness under load.

Why this answer

An ECS capacity provider (or Auto Scaling integration) directly links ECS task-level demand to EC2 instance scaling. When tasks are pending due to insufficient container instances, the capacity provider triggers a scale-out event on the Auto Scaling group, adding EC2 instances to accommodate the workload. This is the most efficient architectural step to reduce placement delays during peak traffic.

Exam trap

The trap here is that candidates may confuse task-level scaling (e.g., Service Auto Scaling) with infrastructure-level scaling, and incorrectly assume that tuning health checks or placement strategies will resolve a capacity shortage caused by insufficient EC2 instances.

Why the other options are wrong

A

Tuning health check settings does not address the root cause of tasks waiting for EC2 capacity; it only affects task stability, not instance availability.

C

Pinning tasks to a single Availability Zone does not address the root cause of insufficient EC2 capacity; it actually reduces fault tolerance and may increase placement constraints, making scaling slower.

D

Switching to Fargate eliminates EC2 scaling concerns but does not address the existing EC2 capacity scaling issue; the question specifically asks for faster scale-out of underlying EC2 capacity, not a migration to a different compute type.

When would these options actually be correct?

A

If the question described tasks frequently failing health checks and being replaced, causing unnecessary scaling events, then tuning health check settings (e.g., increasing grace period or interval) would be the best first step to reduce churn.

C

If the question described a scenario where tasks are failing due to cross-AZ data transfer costs or latency, and the goal is to minimize network overhead, then pinning tasks to a single AZ could be correct.

D

This option would be correct in a scenario where the team wants to eliminate EC2 management entirely and is willing to migrate to serverless compute, such as when the primary goal is to reduce operational overhead and avoid scaling EC2 instances altogether.

Why candidates pick the wrong answer

A

Candidates may confuse task-level health issues with capacity scaling problems, assuming that fixing health checks will reduce the need for new instances.

C

Candidates may think that reducing the number of zones simplifies scheduling and speeds up placement, but they overlook that capacity shortage is the real issue, not placement overhead.

D

Candidates may choose this because Fargate abstracts infrastructure management, making it seem like a simple fix to avoid EC2 scaling problems, without recognizing that the question explicitly asks for a step to improve EC2 scaling, not replace it.

381
MCQeasy

Based on the exhibit, the database must continue serving if the current Availability Zone fails. What should you change?

A.Create a read replica in another Availability Zone and promote it manually if needed.
B.Modify the DB instance to use a Multi-AZ deployment.
C.Increase the automated backup retention period to 30 days.
D.Resize the DB instance to a larger class.
AnswerB

A Multi-AZ RDS deployment provides synchronous standby replication in another Availability Zone and automatic failover if the primary AZ becomes unavailable. This directly matches the requirement to keep the database serving after an AZ failure. It is the simplest resilient design change when the application needs high availability rather than just backups.

Why this answer

Multi-AZ deployment automatically provisions and maintains a synchronous standby replica in a different Availability Zone. If the primary AZ fails, Amazon RDS automatically fails over to the standby, ensuring database availability without manual intervention. This meets the requirement of continuing service during an AZ failure.

Exam trap

The trap here is confusing read replicas (which are for read scaling and asynchronous replication) with Multi-AZ (which is for high availability and synchronous replication), leading candidates to choose Option A for failover scenarios.

Why the other options are wrong

A

Creating a read replica in another AZ does not provide automatic failover; manual promotion introduces downtime, failing the requirement for continuous serving if the current AZ fails.

C

Increasing the automated backup retention period does not provide high availability or automatic failover; it only extends the point-in-time recovery window, which does not help if the current Availability Zone fails.

D

Resizing the DB instance to a larger class improves performance but does not provide failover capability if the current Availability Zone fails. It does not address high availability across AZs.

When would these options actually be correct?

A

This option would be correct if the question required offloading read traffic from the primary database to reduce load, with the ability to promote the replica for disaster recovery, but without a strict requirement for automatic failover.

C

If the question were about meeting a compliance requirement to retain daily backups for 30 days for audit purposes, then increasing the automated backup retention period to 30 days would be the correct answer.

D

This option would be correct in a scenario where the database is experiencing performance bottlenecks due to insufficient compute or memory resources, and the question asks for a solution to improve performance without changing the architecture.

Why candidates pick the wrong answer

A

Candidates may think that a read replica in another AZ provides high availability, but they overlook that failover is manual and not automatic, which is essential for continuous service during an AZ failure.

C

Candidates may confuse backup retention with disaster recovery, thinking that longer backups ensure availability during an AZ failure, but backups require manual restoration and do not provide automatic failover.

D

Candidates may think that a larger instance class inherently provides more reliability or redundancy, confusing vertical scaling with high availability.

382
MCQeasy

A company stores 500 TB of archival data in Amazon S3. The data is accessed only once a year for compliance audits. The company wants the most cost-effective storage solution that still allows retrieval within 48 hours. Which S3 storage class should they use?

A.S3 Intelligent-Tiering
B.S3 Glacier Deep Archive
C.S3 Standard
D.S3 One Zone-IA
AnswerB

S3 Glacier Deep Archive is the lowest-cost storage class, designed for data that is rarely accessed and can tolerate retrieval times of 12 hours or more. It meets the 48-hour retrieval requirement and is ideal for compliance archives. With 500 TB stored and only annual access, this class provides the greatest cost savings compared to other options.

Why this answer

S3 Glacier Deep Archive is the most cost-effective storage class for long-term retention of data that is rarely accessed and can tolerate retrieval times of 12 hours or more. The 48-hour retrieval requirement is satisfied, and the annual access pattern aligns with its design. Other classes either cost more or do not offer the same savings for archival data.

Exam trap

The trap here is choosing S3 One Zone-IA or S3 Intelligent-Tiering because they sound cost-effective, but they are not optimized for data accessed only once a year with a 48-hour retrieval tolerance.

383
MCQeasy

A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most?

A.CloudFront caching with appropriate TTLs
B.AWS Backup Vault Lock
C.IAM Access Analyzer
D.S3 Select
AnswerA

CloudFront caching with appropriate TTLs is correct because it stores static content at edge locations, allowing the website to be served to users from cache even when the S3 origin is temporarily unavailable. As long as the TTL has not expired, users experience no interruption, and the cached content provides a buffer against origin outages. This makes the service more resilient, provided the TTL balances freshness and availability.

Why this answer

CloudFront caches responses from the S3 origin based on configured TTLs (Cache-Control or Expires headers). If the S3 origin experiences a short outage, CloudFront can still serve cached pages to users from its edge locations, maintaining availability. This is the most direct way to ensure users receive content during origin failures.

Exam trap

The trap here is that candidates may confuse CloudFront's caching with other AWS services like S3 Transfer Acceleration or S3 Cross-Region Replication, which do not provide cached responses during origin outages.

How to eliminate wrong answers

Option B (AWS Backup Vault Lock) is wrong because it is a data protection feature for backups, enforcing retention policies and preventing deletion, not related to serving cached web content during origin outages. Option C (IAM Access Analyzer) is wrong because it analyzes resource-based policies to identify unintended public or cross-account access, not for caching or origin failover. Option D (S3 Select) is wrong because it is a feature to retrieve subsets of object data using SQL queries, not for caching or serving static content during outages.

384
MCQmedium

A media company has users around the world uploading 1 to 5 GB files directly to a single Amazon S3 bucket. Upload times are slow from distant regions, but the app must keep using S3 as the destination. What should the architects enable to improve upload performance?

A.Amazon CloudFront for origin caching of uploaded files.
B.Amazon S3 Transfer Acceleration on the bucket.
C.Provisioned IOPS EBS volumes attached to a transfer server.
D.Amazon EFS with a mount target in each Region.
AnswerB

S3 Transfer Acceleration improves upload performance over long distances by routing traffic through AWS edge locations and optimized network paths to the target bucket. This is a strong fit for globally distributed users uploading large files directly to S3. It preserves the same storage destination while making the transfer path faster and more consistent for remote clients.

Why this answer

Amazon S3 Transfer Acceleration (B) uses AWS edge locations to accelerate uploads over the public internet. When a user uploads a file, the data is sent to the nearest edge location via optimized network paths, then forwarded over AWS's private backbone to the S3 bucket. This reduces latency and improves throughput for large files (1–5 GB) from distant regions, directly addressing the slow upload times while keeping S3 as the destination.

Exam trap

The trap here is confusing CloudFront's edge caching for downloads with S3 Transfer Acceleration's edge-based upload optimization, leading candidates to select CloudFront (A) even though it does not improve upload performance to S3.

Why the other options are wrong

A

CloudFront is a content delivery network for caching and accelerating downloads, not uploads. It does not improve upload performance to an S3 bucket because uploads go directly to the origin, not through CloudFront.

C

Provisioned IOPS EBS volumes attached to a transfer server do not improve upload speeds to S3; they improve disk I/O for an intermediate server, but the bottleneck is network latency to S3, not local disk performance.

D

Amazon EFS is a shared file system for EC2 instances, not a direct upload destination for users. The question requires users to upload directly to S3, and EFS cannot replace S3 as the upload target.

When would these options actually be correct?

A

A company needs to reduce latency for global users downloading static content (e.g., images, videos) from an S3 bucket. Enabling CloudFront would cache content at edge locations, improving download speeds and reducing load on the origin.

C

A question where an application requires high-performance, low-latency storage for a database or transactional workload running on a single EC2 instance, and the storage must meet specific IOPS requirements.

D

A company needs a shared file system accessible from multiple EC2 instances across different AWS regions for low-latency file access, with automatic scaling and high durability. Enabling EFS with mount targets in each region would be correct.

Why candidates pick the wrong answer

A

Candidates may assume CloudFront accelerates all data transfer (both upload and download) because it improves delivery speed, but it does not optimize the upload path to S3.

C

Candidates may think that faster local storage on a transfer server will speed up uploads, but the real issue is network distance to S3, not the server's disk speed.

D

Candidates may think that having EFS mount targets in multiple regions would reduce latency for uploads, but they overlook that EFS is not a direct upload endpoint for users and does not replace S3 for object storage.

385
MCQeasy

An application uses DynamoDB to store order status. Reads happen extremely frequently for the same few keys (for example, the most recent orders), and the team wants lower read latency without changing the table’s partition key design. Which AWS service best fits this requirement?

A.Amazon DAX (DynamoDB Accelerator) to cache frequently read items
B.Provision AWS WAF rules to reduce DynamoDB read latency caused by bots
C.Enable multi-region writes in DynamoDB Global Tables to speed up reads locally
D.Add more read capacity units to DynamoDB and avoid caching entirely
AnswerA

DAX is an in-memory caching layer specifically built for DynamoDB. It reduces read latency for hot keys by serving cached responses quickly while still reading from DynamoDB when a key is not cached (or when the cached entry expires). This avoids the need to redesign partition keys.

Why this answer

Amazon DAX (DynamoDB Accelerator) is an in-memory cache that sits between your application and DynamoDB, providing microsecond read latency for frequently accessed items. Because the workload involves extremely frequent reads of the same few keys (hot keys), DAX reduces the load on the DynamoDB table and delivers faster responses without requiring any changes to the partition key design.

Exam trap

The trap here is that candidates often confuse throughput scaling (adding RCUs) with latency optimization, or they mistakenly think Global Tables improve read latency within a single region, when in fact DAX is the only option that directly caches hot keys to reduce read latency without altering the table design.

Why the other options are wrong

B

AWS WAF is a web application firewall that protects against web exploits and bots, but it does not reduce read latency for DynamoDB queries. The latency issue here is due to high read frequency on a few keys, not bot traffic.

C

Multi-region writes in DynamoDB Global Tables improve write availability and disaster recovery, but they do not reduce read latency for frequently accessed hot keys because reads still go to the local region's DynamoDB table without caching. The question specifically asks for lower read latency without changing the partition key design, which DAX addresses via caching.

D

Adding more read capacity units increases throughput but does not reduce latency for hot keys; it only prevents throttling. The question specifically asks for lower read latency, which caching addresses directly.

When would these options actually be correct?

B

A question where an application experiences high read latency due to malicious bot traffic overwhelming DynamoDB, and the requirement is to filter out bot requests before they reach the database, while maintaining low latency for legitimate users.

C

A question asks: 'An application requires low-latency writes and reads across multiple geographic regions, with automatic conflict resolution for the same items written in different regions. Which DynamoDB feature should be used?' In that scenario, DynamoDB Global Tables would be correct.

D

If the question described an application that is experiencing read throttling (ProvisionedThroughputExceededException) on DynamoDB due to insufficient read capacity, and the team wants to eliminate throttling without changing the partition key, then increasing read capacity units would be correct.

Why candidates pick the wrong answer

B

Candidates may think that blocking bots with WAF will reduce overall load and thus latency, but they overlook that the latency issue is caused by legitimate high-frequency reads on hot keys, not by bots.

C

Candidates may think that having a local replica in each region automatically speeds up reads, but Global Tables replicate writes, not reads; reads still hit the local DynamoDB table without caching. The term 'locally' misleads them into believing it reduces read latency.

D

Candidates may think that more capacity automatically means faster reads, overlooking that latency for hot keys is limited by DynamoDB's internal architecture and that caching is the standard solution for sub-millisecond read latency.

386
MCQmedium

A web application for a mobile banking backend is behind an Application Load Balancer. The application must be protected from common SQL injection and cross-site scripting attacks with minimum operational overhead. What should the architect deploy?

A.Security groups on the application instances
B.AWS WAF associated with the Application Load Balancer
C.Network ACLs on the public subnets
D.AWS Shield Advanced only
AnswerB

AWS WAF is a Layer 7 (application-layer) firewall that you can associate with an Application Load Balancer to inspect every HTTP/HTTPS request before it is forwarded to backend instances. It can examine the entire request—including URI, query strings, headers, body, and cookies—and enforce custom rules or AWS managed rule groups specifically designed to block SQL injection and cross-site scripting. This makes it the correct service to protect the mobile banking backend from the described web attacks.

Why this answer

AWS WAF is a web application firewall that helps protect web applications from common web exploits like SQL injection and cross-site scripting (XSS). By associating an AWS WAF web ACL with the Application Load Balancer, you can filter and monitor HTTP(S) requests based on rules that block these attack patterns, all without managing any infrastructure. This provides the required protection with minimal operational overhead because AWS WAF is a fully managed service that integrates directly with ALB.

Exam trap

The trap here is that candidates often confuse network-level controls (security groups, NACLs) with application-layer protection, assuming that blocking ports or IP ranges is sufficient to stop web application attacks like SQL injection and XSS.

How to eliminate wrong answers

Option A is wrong because security groups act as a virtual firewall at the instance level, controlling inbound and outbound traffic based on IP addresses and ports; they cannot inspect application-layer payloads for SQL injection or XSS patterns. Option C is wrong because network ACLs are stateless, operate at the subnet level, and only filter traffic based on IP addresses, ports, and protocols — they have no capability to parse HTTP request bodies or headers for malicious content. Option D is wrong because AWS Shield Advanced provides DDoS protection and cost protection against scaling, but it does not include the application-layer rule sets needed to block SQL injection or XSS attacks; those require a web application firewall like AWS WAF.

387
MCQmedium

Based on the exhibit, a faulty deployment corrupted production data at 10:30 UTC and the issue was discovered at 10:55 UTC. The team needs to recover the database to the last good state before the corruption. Which action should they take?

A.Restore the latest manual snapshot and accept data loss since the snapshot was taken overnight.
B.Use point-in-time restore to create a new database instance at 10:29 UTC, then switch the application to it.
C.Restart the database instance so the transaction log replays the failed migration cleanly.
D.Create a read replica and promote it, because replicas always contain the previous transaction state.
AnswerB

Point-in-time restore is the correct recovery method when automated backups are enabled and the team needs the database just before a known corruption event. Restoring to 10:29 UTC brings the data back to the last safe moment before the migration began. Creating a new instance first avoids modifying the damaged database until the restored copy is validated.

Why this answer

Amazon RDS for MySQL (and other engines) supports point-in-time recovery (PITR), which allows you to restore a database to any second within the backup retention period, up to the last five minutes. By restoring to 10:29 UTC (one minute before the corruption at 10:30 UTC), the team can recover the database to its last good state with minimal data loss. After restoring, the application can be pointed to the new instance, avoiding the corrupted data.

Exam trap

The trap here is that candidates may confuse point-in-time restore with snapshot restore, assuming snapshots are the only recovery option, or incorrectly believe that restarting or promoting a replica can undo a logical corruption that has already been written to disk.

How to eliminate wrong answers

Option A is wrong because restoring the latest manual snapshot would revert the database to the time the snapshot was taken (likely overnight), causing significant data loss of all transactions between that snapshot and 10:30 UTC, which is unacceptable when a more precise recovery is available. Option C is wrong because restarting the database instance does not replay transaction logs to undo a faulty deployment; it only replays committed transactions from the binary logs to ensure consistency, which would reapply the corruption. Option D is wrong because a read replica contains the same data as the primary at the time of replication lag, not a previous transaction state; promoting it would still include the corrupted data if the corruption occurred before the replica caught up.

388
MCQeasy

A travel booking site uses EC2 instances behind an ALB. CPU is consistently high during peak traffic, and request latency rises. What should be configured? The design must avoid adding custom operational scripts.

A.A VPC endpoint for CloudWatch only
B.Auto Scaling policy based on an appropriate CloudWatch metric
C.S3 Object Lock
D.Disable health checks
AnswerB

Auto Scaling adds capacity when load increases and removes it when load falls.

Why this answer

An Auto Scaling policy based on a CloudWatch metric like CPUUtilization or request latency directly addresses the high CPU and rising latency by automatically adding EC2 instances during peak traffic. This eliminates the need for custom scripts and ensures the application scales horizontally to maintain performance.

Exam trap

The trap here is that candidates might think a VPC endpoint (Option A) is needed for CloudWatch metrics, but CloudWatch metrics are already available without a VPC endpoint, and scaling requires an Auto Scaling policy, not just metric access.

How to eliminate wrong answers

Option A is wrong because a VPC endpoint for CloudWatch only enables private connectivity to CloudWatch, but does not provide any scaling or performance improvement for EC2 instances. Option C is wrong because S3 Object Lock is used for data retention and compliance, not for scaling compute resources or reducing latency. Option D is wrong because disabling health checks would cause the ALB to route traffic to unhealthy instances, worsening latency and availability issues.

389
MCQeasy

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The application must be reachable only from a specific corporate CIDR range, and the instances must not be directly reachable from the internet. Which combination of security group configurations meets these requirements?

A.Allow inbound HTTP and HTTPS from 0.0.0.0/0 on the load balancer security group, and allow inbound HTTP from the load balancer security group on the instance security group.
B.Allow inbound HTTP and HTTPS from the corporate CIDR range on the instance security group, and allow outbound traffic to the load balancer security group.
C.Allow inbound HTTP and HTTPS from the corporate CIDR range on both the load balancer security group and the instance security group.
D.Allow inbound HTTP and HTTPS from the corporate CIDR range on the load balancer security group, and allow inbound HTTP from the load balancer security group on the instance security group.
AnswerD

Restricting the load balancer security group to the corporate CIDR range limits access to the approved network. Referencing the load balancer security group as the source in the instance security group allows only traffic forwarded by the load balancer, so the instances are not directly reachable from the internet. This matches both requirements.

Why this answer

The load balancer security group should allow only the corporate CIDR range on the listener ports, and the instance security group should allow traffic only from the load balancer security group. Referencing a security group as a source is the cleanest way to allow traffic from the load balancer without hardcoding its IP addresses, and it prevents direct client access to the instances.

Exam trap

The trap here is putting the corporate CIDR range on the instance security group, which permits clients to bypass the load balancer rather than forcing all traffic through it.

390
MCQeasy

A system uses multiple AWS Lambda functions behind different event sources. One Lambda occasionally spikes and causes other Lambdas to be throttled due to shared concurrency limits. Which setting best helps ensure the important Lambda keeps capacity during spikes?

A.Increase the function timeout so throttling is less likely.
B.Set Reserved Concurrency for the important Lambda function.
C.Enable Provisioned Concurrency for every Lambda in the account.
D.Reduce the number of IAM policies attached to the Lambda roles.
AnswerB

Reserved concurrency allocates a guaranteed amount of concurrent execution capacity to a specific Lambda. This prevents other functions from consuming all concurrency and throttling the important one. If the reserved limit is reached, only that function is throttled, isolating impact.

Why this answer

Reserved Concurrency guarantees a set number of concurrent executions for a specific Lambda function, isolating it from the account-level concurrency pool. This ensures that the important function always has capacity available, even when other functions spike and consume the shared pool. Without this setting, all functions compete for the same 1,000 concurrent executions (default regional limit), and a spike in one can throttle others.

Exam trap

The trap here is that candidates confuse Provisioned Concurrency (which reduces cold starts) with Reserved Concurrency (which guarantees capacity), leading them to choose Option C, even though Provisioned Concurrency does not protect against throttling from other functions.

Why the other options are wrong

A

Increasing the function timeout does not affect concurrency limits; it only extends the maximum execution duration. Throttling occurs due to hitting the account-level concurrency limit, not timeout.

C

Provisioned Concurrency for every Lambda would waste resources and cost, and it does not guarantee that the important Lambda has reserved capacity during spikes; it only pre-warms instances, but all functions still share the same regional concurrency limit.

D

Reducing IAM policies does not affect Lambda concurrency limits; it only reduces permissions, which does not prevent throttling from shared concurrency.

When would these options actually be correct?

A

A question where a Lambda function is timing out before completing its task, causing errors or retries, and the requirement is to ensure it finishes processing without failure.

C

A question where the goal is to reduce cold starts for all functions in a latency-sensitive application, and cost is not a primary concern. For example: 'A company wants to ensure all its Lambda functions have minimal latency during unpredictable traffic spikes. Which action should be taken?'

D

In a scenario where a Lambda function is failing due to permission errors (e.g., AccessDeniedException) and you need to minimize the attack surface, reducing unnecessary IAM policies would be correct to follow the principle of least privilege.

Why candidates pick the wrong answer

A

Candidates may confuse timeout with concurrency, thinking longer execution reduces the chance of being throttled, or believe that throttling is related to execution duration.

C

Candidates may think Provisioned Concurrency reserves capacity, but it only pre-initializes instances; it does not prevent other functions from consuming the shared pool. The term 'Provisioned' sounds like reserving capacity, leading to confusion with Reserved Concurrency.

D

Candidates may mistakenly think that reducing IAM policies reduces resource usage or overhead, thereby preventing throttling, but IAM policies have no impact on concurrency limits.

391
MCQhard

A DynamoDB table for a retail API has a partition key based only on the current date. Write throttling occurs during business hours. What is the best design change? The design must avoid adding custom operational scripts.

A.Use a higher-cardinality partition key that distributes writes across partitions
B.Create a global secondary index with the same date key
C.Reduce the table's write capacity
D.Move the table to S3 Glacier Instant Retrieval
AnswerA

A partition key with only two distinct values funnels all writes into two partitions, each capped at 1,000 write capacity units, so a high write volume quickly throttles. Introducing a higher-cardinality key—such as customer ID or a composite key with a random suffix—spreads writes across many partitions, letting the table utilize its full provisioned capacity and avoid any single hot partition.

Why this answer

Using only the current date as a partition key creates a 'hot partition' because all writes for the day target a single partition, exceeding its 1,000 WCU limit. A higher-cardinality partition key (e.g., combining date with user ID or order ID) distributes writes evenly across partitions, eliminating throttling without custom scripts.

Exam trap

The trap here is that candidates often confuse GSIs as a solution for write hot spots, but GSIs only help with read patterns and do not change the base table's write distribution.

How to eliminate wrong answers

Option B is wrong because a global secondary index (GSI) inherits the same write capacity from the base table and does not redistribute the write load; it would still be throttled. Option C is wrong because reducing write capacity would worsen throttling, not solve it. Option D is wrong because S3 Glacier Instant Retrieval is for archival data with infrequent access, not for a DynamoDB table requiring low-latency writes for a retail API.

392
MCQmedium

A media processing pipeline uses EBS-backed storage for an application that performs sustained random I/O with low latency requirements. During peak processing windows, the team sees increased read latency and occasional timeouts at the application layer. They need predictable, high IOPS performance rather than best-effort throughput. Which EBS configuration choice is most appropriate?

A.Use gp2 volumes and rely on burst credits to handle peak random I/O latency requirements.
B.Use io1 or io2 EBS volumes configured with a high provisioned IOPS value, and attach them to EBS-optimized instances.
C.Use standard HDD (st1) volumes, because they provide high throughput and will reduce latency automatically.
D.Use S3 instead of EBS for random I/O latency reduction without changing the application.
AnswerB

io1/io2 are designed for predictable, low-latency IOPS for sustained I/O workloads. By provisioning a sufficient IOPS level, you improve consistency during peak windows. Using EBS-optimized instances ensures the instance-to-EBS bandwidth and I/O performance are adequate so the instance does not become the bottleneck before EBS can deliver the provisioned IOPS.

Why this answer

Io1 and io2 volumes are provisioned IOPS SSD volumes designed for sustained, predictable high IOPS performance, which directly addresses the application's need for low-latency random I/O during peak loads. Attaching them to EBS-optimized instances ensures dedicated network bandwidth for EBS traffic, eliminating contention and preventing timeouts.

Exam trap

The trap here is that candidates may choose gp2 (Option A) assuming burst credits will cover peak loads, but they fail to recognize that sustained peak I/O exhausts credits, leading to performance degradation, whereas provisioned IOPS volumes guarantee consistent performance regardless of duration.

How to eliminate wrong answers

Option A is wrong because gp2 volumes rely on burst credits that can be exhausted during sustained peak I/O, leading to throttled performance and increased latency, not predictable high IOPS. Option C is wrong because st1 volumes are HDD-based and optimized for sequential throughput, not random I/O; they cannot provide low latency or high IOPS for random access patterns. Option D is wrong because S3 is an object storage service with higher latency and no support for low-latency random I/O; it cannot replace EBS for block-level access without significant application changes.

393
MCQhard

A SaaS vendor’s automation account in Account B needs to assume a role in a customer account in Account A to read a specific S3 bucket and publish a deployment status file. The customer is worried about confused deputy attacks because multiple customers use the same vendor software. Which trust-policy design best meets the requirement?

A.Allow the Account B root principal to assume the role if the caller knows the role ARN.
B.Allow only the vendor’s specific IAM principal to assume the role and require a unique sts:ExternalId condition.
C.Attach a permissions boundary to the role so that the vendor cannot exceed the approved permissions.
D.Require MFA for the role assumption because it ensures only the vendor’s production automation can use the role.
AnswerB

This is the standard confused deputy protection pattern for third-party cross-account access. The trust policy limits who can call AssumeRole, and the sts:ExternalId condition lets the customer require a customer-specific value that the vendor must supply. That prevents another customer or a malicious party from reusing the same role ARN successfully.

Why this answer

The `sts:ExternalId` condition is specifically designed to prevent the confused deputy problem in cross-account role assumptions. By requiring a unique external ID that only the customer knows, the customer ensures that the vendor's automation can only assume the role when acting on behalf of that specific customer, even if multiple customers use the same vendor software.

Exam trap

The trap here is that candidates often confuse MFA or permissions boundaries as solutions for the confused deputy problem, when in fact only the `sts:ExternalId` condition directly mitigates this specific threat by providing a customer-specific identifier in the trust policy.

How to eliminate wrong answers

Option A is wrong because allowing the Account B root principal to assume the role based solely on knowing the role ARN provides no protection against confused deputy attacks; any IAM entity in Account B (including compromised or malicious principals) could assume the role. Option C is wrong because a permissions boundary limits the maximum permissions the role can grant, but it does not address the confused deputy threat; it controls scope, not identity verification. Option D is wrong because requiring MFA ensures the caller is authenticated via a second factor, but it does not prevent a confused deputy scenario where the vendor's automation could be tricked into assuming the role on behalf of a different customer; MFA does not provide a customer-specific identifier.

394
MCQeasy

An order system receives events and uses a Lambda function to write each order into a database. During traffic spikes, the database sometimes throttles, and Lambda retries lead to occasional message loss in the event flow. The team wants buffering, automatic retries, and a way to isolate messages that repeatedly fail so they can be inspected later. What design change best meets this need?

A.Send events directly from EventBridge to Lambda without any queue to simplify the flow.
B.Use Amazon SQS as a buffer between the event source and Lambda, with an SQS dead-letter queue (DLQ).
C.Use SNS fan-out to multiple Lambda functions, but keep no retry logic and no DLQ.
D.Store events in an S3 bucket and trigger Lambda immediately after each upload, without using DLQs.
AnswerB

Using SQS as a buffer between the event source and Lambda is correct because SQS decouples event producers from the consumer, smoothing traffic spikes by buffering messages until Lambda can poll them. The visibility timeout provides built-in retry logic: if Lambda fails to process a message, it becomes visible again for another attempt, and after a configured Maximum Receives threshold, the message is automatically diverted to a dead-letter queue for later analysis. This pattern ensures that no order event is lost, supports backpressure, and lets you isolate recurring processing failures without disrupting the continuous flow of valid events.

Why this answer

Amazon SQS acts as a durable buffer between the event source and Lambda, absorbing traffic spikes and decoupling the producer from the consumer. The SQS dead-letter queue (DLQ) automatically captures messages that exceed the configured maximum retries, allowing the team to inspect and reprocess them later without loss. This design provides the required buffering, automatic retries via the Lambda event source mapping, and isolation of repeatedly failing messages.

Exam trap

The trap here is that candidates often assume a direct event-driven flow (like EventBridge to Lambda) is simpler and sufficient, but they overlook the need for buffering and a DLQ to handle throttling and isolate persistent failures, which SQS explicitly provides.

Why the other options are wrong

A

Sending events directly from EventBridge to Lambda without a queue provides no buffering during traffic spikes, so Lambda retries still cause throttling and message loss. It also lacks a dead-letter queue to isolate repeatedly failing messages.

C

SNS fan-out without retry logic or a DLQ does not provide buffering, automatic retries, or isolation of failed messages, which are explicitly required to handle throttling and prevent message loss.

D

S3 event notifications have no built-in retry logic or dead-letter queue; if Lambda fails, the event is lost after the retry limit, failing to isolate repeatedly failing messages for inspection.

When would these options actually be correct?

A

A question where the requirement is to minimize latency and cost for a low-volume, predictable event flow, and where message loss is acceptable. For example: 'A logging system processes non-critical events that can tolerate occasional drops; the team wants the simplest architecture.'

C

A system needs to broadcast the same event to multiple independent downstream services (e.g., notification, analytics, audit) simultaneously, and each service has its own retry and error handling. SNS fan-out is ideal for this parallel distribution.

D

For a batch processing system where raw data must be durably stored for compliance and reprocessing, and failures are handled by a separate monitoring process, S3 triggering Lambda immediately is appropriate.

Why candidates pick the wrong answer

A

Candidates may think removing the queue simplifies the architecture and reduces latency, overlooking the need for buffering and retry isolation during spikes.

C

Candidates may think SNS fan-out adds reliability through multiple subscribers, but they overlook that without retries and a DLQ, failed messages are lost, failing to meet the question's requirements for buffering and error isolation.

D

Candidates may think S3 provides durable storage and automatic event triggers, overlooking the lack of retry and DLQ capabilities needed for handling transient failures and isolating poison messages.

395
MCQmedium

A company uses Amazon RDS for a PostgreSQL database powering a customer-facing application. The application’s availability depends on fast database failover with minimal manual intervention. The RDS instance currently runs as a single-AZ deployment in one DB subnet group. Which change most directly meets the goal?

A.Create a read replica in a different Availability Zone and configure the application to fail over manually.
B.Enable Multi-AZ for the RDS DB instance so AWS manages a standby in another Availability Zone with automatic failover.
C.Switch the database to use EBS snapshots more frequently and restore in case of failure.
D.Pin the DB to a specific instance type with higher CPU credits to prevent CPU-related disconnects.
AnswerB

RDS Multi-AZ creates a synchronous standby replica in a separate Availability Zone and automatically flips the DNS endpoint to the standby within seconds if the primary instance fails, a network partition occurs, or an AZ becomes unavailable. This is the only option that provides true high availability with no manual intervention, making it the correct answer because it directly addresses resilience to an entire Availability Zone failure.

Why this answer

Enabling Multi-AZ for the RDS DB instance creates a synchronous standby replica in a different Availability Zone. AWS automatically handles failover to the standby with no manual intervention required, which directly meets the goal of fast database failover with minimal manual intervention.

Exam trap

The trap here is that candidates often confuse read replicas (which are for read scaling and manual promotion) with Multi-AZ (which provides automatic failover and high availability), leading them to choose Option A incorrectly.

Why the other options are wrong

A

This option requires manual failover, which contradicts the requirement for minimal manual intervention and fast failover.

C

Frequent EBS snapshots provide point-in-time recovery but do not enable fast, automatic failover; restoring from a snapshot requires manual intervention and takes significant time, failing to meet the minimal manual intervention and fast failover requirements.

D

Higher CPU credits prevent CPU-related throttling but do not address database failover speed or minimize manual intervention, which is the core requirement for fast failover in a single-AZ deployment.

When would these options actually be correct?

A

A question where the application can tolerate some downtime but needs to offload read traffic, and the exam asks for a cost-effective way to improve read scalability without automatic failover.

C

A question where the goal is to minimize data loss (low RPO) for a non-critical application that can tolerate downtime for recovery, and the requirement is to implement a cost-effective backup strategy without needing automatic failover.

D

For a question where the application experiences intermittent performance degradation due to CPU credit exhaustion on a burstable instance (e.g., t3/t4g), and the goal is to eliminate CPU-related throttling without changing architecture.

Why candidates pick the wrong answer

A

Candidates may think a read replica in a different AZ provides high availability, but they overlook that failover is manual and not automatic.

C

Candidates may confuse backup and restore (snapshots) with high availability, thinking that more frequent backups can serve as a failover mechanism, or they underestimate the time and manual effort required to restore from a snapshot.

D

Candidates may mistakenly think that preventing CPU issues is equivalent to ensuring high availability, or they may confuse performance tuning with failover mechanisms.

396
MCQmedium

A security analyst needs to let an external vendor (AWS account 555566667777) read data from a set of internal resources in your AWS account. You created an IAM role called VendorReadRole with a policy that allows the required API calls. However, when the vendor tries to access, CloudTrail shows the call fails at AssumeRole with: "Not authorized to perform: sts:AssumeRole". What is the most appropriate fix?

A.Add an allow statement for the vendor in the role’s trust policy to permit sts:AssumeRole from the vendor account (and include any required ExternalId condition).
B.Attach the same allow policy to the vendor account’s existing IAM user so the user can call sts:AssumeRole directly into your role.
C.Replace the AssumeRole call with GetCallerIdentity so the vendor can infer permissions without assuming the role.
D.Enable MFA on the vendor’s IAM user and require MFA for your role using condition keys in the permissions policy.
AnswerA

To grant an external vendor access to your AWS resources, you must edit the role's trust policy to include a principal from the vendor account and an action of sts:AssumeRole. This trust relationship is the only mechanism that authorizes a foreign principal to assume your role; the permissions policy alone cannot authorize cross-account assumption. Adding an ExternalId condition prevents the confused deputy problem by ensuring the role is assumed only for your intended vendor, not a third party using the same role.

Why this answer

The error 'Not authorized to perform: sts:AssumeRole' indicates that the role's trust policy does not grant the external AWS account (555566667777) permission to assume the role. The trust policy must include an Allow statement with the sts:AssumeRole action, specifying the external account as the principal, and optionally an ExternalId condition to prevent the confused deputy problem. Without this trust policy configuration, even if the permissions policy allows the required API calls, the vendor cannot assume the role.

Exam trap

The trap here is that candidates often confuse the role's permissions policy (which defines what actions the role can perform) with the trust policy (which defines who can assume the role), leading them to incorrectly modify the permissions policy or the vendor's IAM user instead of the trust policy.

Why the other options are wrong

B

The trust policy on the IAM role must explicitly allow the external account to assume the role; attaching a policy to the vendor's IAM user does not grant cross-account sts:AssumeRole permissions because the role's trust policy controls who can assume it.

C

GetCallerIdentity does not grant cross-account access; it only returns details about the caller's own identity. The vendor needs to assume a role to access resources in the other account, not just check who they are.

D

The error is 'Not authorized to perform: sts:AssumeRole', which is a trust policy issue, not a permissions policy issue. Enabling MFA on the vendor's IAM user and adding an MFA condition to the role's permissions policy does not grant the vendor permission to assume the role; the trust policy must explicitly allow the vendor account to call sts:AssumeRole.

When would these options actually be correct?

B

If the vendor needed to access resources in your account using their own IAM user (not assuming a role), you would attach a resource-based policy (e.g., S3 bucket policy) allowing the vendor's user ARN. In that case, the vendor's user policy would need to allow the corresponding API calls (e.g., s3:GetObject).

C

When troubleshooting an IAM permissions issue within the same account, GetCallerIdentity can be used to verify the identity and permissions of the caller without needing to assume a role.

D

A question where the vendor is already allowed to assume the role (trust policy is correct) but the role's permissions policy restricts access unless MFA is present. For example, a security requirement that all cross-account role assumptions must be authenticated with MFA. In that case, adding an MFA condition to the permissions policy would be the correct fix.

Why candidates pick the wrong answer

B

Candidates may think that adding a policy to the vendor's IAM user is sufficient to allow cross-account access, but they overlook that the role's trust policy is the gatekeeper for sts:AssumeRole.

C

Candidates may confuse GetCallerIdentity with a method to gain permissions, thinking it can infer or grant access, when it is only a diagnostic API.

D

Candidates may think that adding MFA strengthens security and thus solves authorization issues, or they confuse the trust policy (who can assume the role) with the permissions policy (what actions the role can perform). They might also believe that MFA is a universal solution for access denied errors.

397
MCQmedium

A public API for a e-learning platform is deployed on API Gateway. Clients must authenticate with standards-based tokens issued by an external OpenID Connect provider. Which authorization mechanism should be used?

A.A VPC endpoint policy
B.IAM authorization for all internet users
C.API keys only
D.JWT authorizer configured for the OpenID Connect issuer
AnswerD

The JWT authorizer configured for the OpenID Connect issuer is the correct choice because it validates the JWT's signature, expiry, issuer, and audience against the OIDC provider's public keys (JWKS), requiring no custom Lambda or additional infrastructure. It integrates directly with any standards-compliant IdP—such as Auth0, Okta, or the platform's existing identity service—so users' existing login tokens are recognized. This approach authenticates every request with low latency and minimal operational overhead while supporting fine-grained scopes and claims for authorization.

Why this answer

API Gateway supports JWT authorizers that validate JSON Web Tokens (JWTs) issued by an external OpenID Connect (OIDC) provider. This allows the API to authenticate clients using standards-based tokens without managing a custom Lambda authorizer, and it directly integrates with the OIDC issuer's JWKS endpoint to verify token signatures.

Exam trap

The trap here is that candidates often confuse API keys (which only identify the caller for usage plans) with authentication mechanisms, or assume IAM authorization can validate third-party OIDC tokens, when in fact IAM authorization requires AWS credentials, not external tokens.

How to eliminate wrong answers

Option A is wrong because a VPC endpoint policy controls access to API Gateway from within a VPC, not authentication for internet-based clients using OIDC tokens. Option B is wrong because IAM authorization is designed for AWS-authenticated principals (e.g., IAM users, roles) and does not validate tokens from external OpenID Connect providers; it uses AWS Signature Version 4, not OIDC tokens. Option C is wrong because API keys only provide simple rate limiting and usage plans, not authentication or authorization; they do not validate the identity of the caller or support OIDC token verification.

398
MCQmedium

A media archive requires consistent high IOPS for a transactional database on EC2. Which EBS volume type is most suitable?

A.Provisioned IOPS SSD such as io2
B.st1 Throughput Optimized HDD
C.Instance store only
D.sc1 Cold HDD
AnswerA

Provisioned IOPS SSD volumes such as io2 are engineered for latency-sensitive, transaction-heavy workloads that require consistent and predictable performance. They let you specify a guaranteed IOPS level independent of volume size, with up to 64,000 IOPS per volume when using supported instances, and offer 99.999% durability. This makes io2 ideal for a media archive that demands both high random I/O throughput and reliable, low-latency access.

Why this answer

The scenario requires consistent high IOPS for a transactional database, which demands low-latency, predictable performance. Provisioned IOPS SSD volumes like io2 are designed specifically for such workloads, offering up to 256,000 IOPS per volume with 99.999% durability, making them the most suitable choice for consistent high IOPS.

Exam trap

The trap here is that candidates often confuse throughput-optimized HDD (st1) with IOPS-focused workloads, mistakenly thinking high throughput equals high IOPS, but IOPS measures random access operations while throughput measures sequential data transfer, and transactional databases require low-latency random I/O.

How to eliminate wrong answers

Option B (st1 Throughput Optimized HDD) is wrong because it is a throughput-optimized HDD volume designed for large, sequential workloads like big data and log processing, not for transactional databases requiring consistent high IOPS and low latency. Option C (Instance store only) is wrong because instance store volumes provide ephemeral storage that is not persistent; data is lost if the instance stops or terminates, making it unsuitable for a transactional database that requires data durability. Option D (sc1 Cold HDD) is wrong because it is a cold HDD volume optimized for infrequently accessed data with the lowest cost, offering very low IOPS and throughput, which cannot meet the consistent high IOPS demands of a transactional database.

399
MCQmedium

A high-volume telemetry pipeline writes streaming click events that must be processed by multiple independent consumers. Which service is most appropriate? The design must avoid adding custom operational scripts.

A.Amazon Kinesis Data Streams
B.AWS DataSync
C.Amazon EBS
D.Amazon Route 53
AnswerA

Amazon Kinesis Data Streams is a massively scalable, real-time data streaming service designed for high-throughput ingestion of click events and telemetry. Data is written to shards as durable records, allowing multiple consumer applications (e.g., via the Kinesis Client Library, Lambda, or enhanced fan-out) to read and process the same stream independently. It also supports replay and configurable retention up to 365 days, making it ideal for pipeline architectures with multiple downstream consumers.

Why this answer

Amazon Kinesis Data Streams is designed for real-time streaming of high-volume data, such as click events, and allows multiple independent consumers to process the same stream concurrently via enhanced fan-out or shared throughput. It provides durable, ordered data retention and integrates with AWS Lambda, Kinesis Data Analytics, and Kinesis Data Firehose without requiring custom operational scripts.

Exam trap

The trap here is that candidates may confuse AWS DataSync or EBS as viable for streaming data, but DataSync is for batch file transfers and EBS is for block storage, neither supporting real-time, multi-consumer event processing.

How to eliminate wrong answers

Option B (AWS DataSync) is wrong because it is a data transfer service for moving large datasets between on-premises storage and AWS services (e.g., S3, EFS) over the internet or Direct Connect, not a real-time streaming pipeline for multiple consumers. Option C (Amazon EBS) is wrong because it provides block-level storage volumes for EC2 instances, not a streaming data ingestion or processing service, and cannot support multiple independent consumers reading the same event stream. Option D (Amazon Route 53) is wrong because it is a DNS web service for domain name resolution and traffic routing, not a data streaming or processing service.

400
MCQmedium

A financial analytics firm runs a nightly batch job on a fleet of Amazon EC2 instances that read millions of small JSON objects from an Amazon S3 bucket and write aggregated results to another S3 bucket. The job currently takes over six hours and the team wants to reduce this time without modifying the application code. The S3 buckets are in the same AWS Region as the EC2 instances. Which action will most effectively improve the performance of the batch job?

A.Increase the number of EC2 instances and ensure the S3 keys are distributed across multiple prefixes.
B.Switch the batch job to use S3 Select to retrieve only the necessary fields from each JSON object.
C.Configure the EC2 instances to use S3 Transfer Acceleration for all requests.
D.Enable S3 Cross-Region Replication to a bucket in a different Region and redirect the job to the replica.
AnswerA

S3 automatically scales to high request rates, but the achieved throughput depends on the number of prefixes and the parallelism of requests. By adding more EC2 instances and spreading object keys across many prefixes, the job can issue more concurrent GET and PUT requests, each prefix supporting thousands of requests per second. This directly reduces the overall processing time without code changes.

Why this answer

The batch job's performance is constrained by the rate at which it can read and write objects to S3. S3 scales horizontally, but the achievable request rate depends on the number of distinct key prefixes and the level of concurrency. By adding more EC2 instances and distributing keys across multiple prefixes, the job can parallelize requests and fully utilize S3's scalability, dramatically reducing the six-hour runtime without modifying application logic.

Exam trap

The trap here is assuming that S3 Transfer Acceleration or Cross-Region Replication will speed up same-Region access, when the real bottleneck is per-prefix request rate and parallelism.

401
Multi-Selectmedium

A media company is designing a high-performance architecture to serve video content to users worldwide. The solution must minimize latency for end users and reduce the load on the origin servers. The video files are stored in an Amazon S3 bucket. Which three options should be combined to meet these requirements? (Choose three.)

Select 3 answers
.Use Amazon CloudFront as a content delivery network (CDN) with the S3 bucket as the origin.
.Enable S3 Transfer Acceleration on the bucket to speed up uploads.
.Configure CloudFront to use Regional Edge Caches to improve cache hit ratios for less popular content.
.Use Amazon ElastiCache for Memcached to cache video metadata at the edge.
.Enable S3 default encryption using AWS KMS to improve data transfer performance.
.Implement origin shield in CloudFront to reduce the number of requests sent to the S3 origin.

Why this answer

Amazon CloudFront as a CDN with the S3 bucket as the origin minimizes latency by caching video content at edge locations worldwide, serving users from the nearest edge. This reduces load on the origin S3 bucket by handling requests at the edge. Regional Edge Caches further improve cache hit ratios for less popular content by caching it at regional locations, reducing the need to fetch from the origin.

Origin shield in CloudFront consolidates requests from multiple edge locations into a single request to the S3 origin, significantly reducing the number of direct requests and lowering origin load.

Exam trap

The trap here is that candidates may confuse S3 Transfer Acceleration (which optimizes uploads) with CloudFront (which optimizes downloads), or think that ElastiCache can be used as a CDN for video content, when it is actually an in-memory cache for application data, not for serving static files at the edge.

402
MCQmedium

A Lambda function for a healthcare document service needs to read a database password. The password must rotate automatically every 30 days and should not be stored in environment variables. Which service should be used?

A.A KMS-encrypted Lambda environment variable
B.An encrypted object in Amazon S3
C.AWS Systems Manager Parameter Store SecureString without automation
D.AWS Secrets Manager with rotation enabled
AnswerD

AWS Secrets Manager is purpose-built for storing secrets and, when rotation is enabled, it automatically rotates the secret value on a schedule using an associated Lambda function. It also keeps previous versions during rotation so applications can continue to work while the new secret is being tested, and it integrates natively with services like RDS, Redshift, and DocumentDB. This directly satisfies the need for automated secret rotation and eliminates the operational overhead of manually updating credentials.

Why this answer

AWS Secrets Manager is the correct choice because it is designed specifically for storing and automatically rotating database credentials. It supports native rotation for Amazon RDS, Redshift, and DocumentDB with a built-in Lambda rotation function, and it can rotate secrets on a schedule (e.g., every 30 days) without storing the password in environment variables. This meets the healthcare document service's requirement for automatic rotation and secure storage.

Exam trap

The trap here is that candidates often confuse Systems Manager Parameter Store (which can store SecureStrings) with Secrets Manager, but Parameter Store lacks automatic rotation, making it unsuitable for a 30-day rotation requirement without additional custom automation.

How to eliminate wrong answers

Option A is wrong because storing a KMS-encrypted password in a Lambda environment variable does not support automatic rotation; you would have to manually update the environment variable and redeploy the function. Option B is wrong because an encrypted object in Amazon S3 is a static storage mechanism with no built-in rotation capability, and accessing it requires managing S3 permissions and decryption logic manually. Option C is wrong because AWS Systems Manager Parameter Store SecureString without automation can store a secure password but lacks native rotation scheduling; you would need to build a custom rotation solution, whereas Secrets Manager provides this out of the box.

403
Multi-Selecthard

A serverless checkout API uses AWS Lambda behind API Gateway. Every weekday at 09:00 UTC, marketing triggers a predictable surge. The first few minutes after each surge show cold-start latency, but traffic volume is forecastable and the business wants stable p95 latency. Which two changes should the team implement? Select two.

Select 2 answers
A.Publish a Lambda version and attach provisioned concurrency to an alias that points to that version.
B.Use Application Auto Scaling scheduled actions to raise provisioned concurrency before 09:00 UTC and lower it afterward.
C.Increase the Lambda timeout so the function has more time to initialize during the spike.
D.Double the memory size during the spike without changing the concurrency model.
E.Move the function into more Availability Zones so the platform can spread cold starts across regions.
AnswersA, B

Provisioned concurrency keeps execution environments initialized and ready to serve requests, which is the correct way to reduce cold starts. Using an alias tied to a published version is the standard deployment pattern for managing that setting safely. This directly improves p95 latency during predictable bursts.

Why this answer

Provisioned concurrency keeps a specified number of Lambda execution environments initialized and ready to respond immediately, eliminating cold starts for predictable traffic patterns. By publishing a Lambda version and attaching provisioned concurrency to an alias pointing to that version, the team ensures that the surge at 09:00 UTC is handled without cold-start latency, stabilizing p95 latency.

Exam trap

The trap here is that candidates often confuse increasing Lambda timeout or memory with solving cold-start latency, but these settings do not pre-warm execution environments; only provisioned concurrency (and optionally scheduled scaling) directly eliminates cold starts for predictable surges.

404
MCQeasy

A company serves a public API through a CloudFront distribution. They want to automatically block common web exploits (for example, OWASP Top 10–style threats) without building custom detection logic. Which AWS service configuration best meets the goal?

A.Enable AWS WAF with AWS Managed Rules and associate the web ACL with the CloudFront distribution.
B.Enable AWS Shield Advanced only; it fully replaces the need for WAF rule evaluation.
C.Attach a security group rule to the ALB to block malicious patterns based on HTTP request bodies.
D.Use Security Hub to block requests automatically when it detects suspicious activity.
AnswerA

AWS WAF inspects HTTP(S) requests and applies allow/block decisions based on rule matches. AWS Managed Rules provide prebuilt protections for common threat patterns, and attaching the WAF web ACL to CloudFront applies filtering at the edge.

Why this answer

AWS WAF with AWS Managed Rules provides pre-configured rule sets specifically designed to block common web exploits, including OWASP Top 10 threats, without requiring custom detection logic. By associating the web ACL with a CloudFront distribution, the filtering occurs at the edge, protecting the origin from malicious traffic before it reaches the application.

Exam trap

The trap here is confusing AWS Shield Advanced (which handles volumetric DDoS attacks) with AWS WAF (which handles application-layer threats like OWASP Top 10), leading candidates to believe Shield alone can replace WAF rule evaluation.

How to eliminate wrong answers

Option B is wrong because AWS Shield Advanced provides DDoS protection and cost mitigation, but it does not include application-layer rule evaluation for OWASP Top 10 threats; it is not a replacement for WAF. Option C is wrong because security groups operate at the network layer (Layer 3/4) and cannot inspect HTTP request bodies or application-layer payloads to block patterns like SQL injection or XSS. Option D is wrong because AWS Security Hub is a security posture management service that aggregates findings and does not have the capability to automatically block requests in real-time; it lacks inline traffic inspection and enforcement actions.

405
MCQmedium

A company runs a media-processing pipeline that ingests thousands of small files per minute into Amazon S3 and triggers AWS Lambda functions for each object. Processing each file takes 2-3 seconds, and the team is seeing throttling errors and duplicated processing under load. They want to decouple ingestion from processing, buffer bursty traffic, and avoid duplicate deliveries to Lambda. Which solution should a solutions architect recommend?

A.Configure an Amazon S3 event notification to send events to an Amazon SQS standard queue, then have Lambda poll the queue with a batch size of 10 and a visibility timeout longer than the processing time.
B.Increase the Lambda reserved concurrency to the maximum account limit and keep the direct S3 event notification to Lambda, so functions run in parallel and finish faster.
C.Use an S3 event notification to invoke an AWS Step Functions state machine that processes each file and retries on failure, with a Map state for parallelism.
D.Enable S3 Transfer Acceleration on the bucket and configure Lambda provisioned concurrency to handle the burst of events without cold starts.
AnswerA

S3 event notifications to SQS decouple ingestion from processing, and Lambda polling the queue scales with the backlog. A standard queue provides high throughput; setting the visibility timeout longer than the function runtime prevents duplicate processing because the message stays invisible until the function completes. This design absorbs bursts and avoids Lambda throttling caused by direct invocation.

Why this answer

Decoupling ingestion from processing with an SQS queue lets the pipeline absorb bursts while Lambda scales consumers based on backlog depth. A visibility timeout longer than the function execution time ensures that a message is not redelivered before processing completes, which prevents duplicates. Direct invocation from S3 does not provide buffering, and concurrency or acceleration features do not solve throttling or duplicate delivery.

Exam trap

The trap here is assuming that increasing Lambda concurrency or adding provisioned concurrency resolves throttling from bursty direct event sources, when the real fix is to insert a buffer such as Amazon SQS between the producer and the consumer.

406
MCQmedium

A trading dashboard runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include?

A.A single EC2 instance with detailed monitoring
B.Subnets in at least two Availability Zones with health checks enabled
C.All instances in one larger subnet
D.A Network Load Balancer in one subnet
AnswerB

Placing subnets in at least two Availability Zones (AZs) and attaching health checks to the Auto Scaling group allows the group to detect and replace unhealthy instances while maintaining desired capacity across AZs. If one AZ fails, the remaining healthy instances in the other AZ continue serving traffic, and Auto Scaling launches new instances in the surviving AZs to compensate. Health checks (ELB or EC2 status checks) drive the replacement process, and distributing subnets across AZs makes the architecture resilient to both instance-level and AZ-level failures, which is the key requirement for a fault-tolerant trading dashboard.

Why this answer

Distributing EC2 instances across subnets in at least two Availability Zones ensures that if one AZ fails, the Auto Scaling group can maintain capacity using instances in the remaining AZ(s). Enabling health checks allows the group to detect and replace unhealthy instances, which is essential for fault tolerance. This configuration meets the requirement to tolerate the failure of one Availability Zone.

Exam trap

The trap here is that candidates often confuse high availability with fault tolerance, thinking a single large subnet or a single instance with monitoring is sufficient, when in fact distributing across multiple Availability Zones is the key to surviving an AZ failure.

How to eliminate wrong answers

Option A is wrong because a single EC2 instance, even with detailed monitoring, cannot tolerate the failure of an entire Availability Zone; if that AZ goes down, the instance becomes unavailable. Option C is wrong because placing all instances in one larger subnet confines them to a single Availability Zone, providing no redundancy if that AZ fails. Option D is wrong because a Network Load Balancer in one subnet does not solve the AZ failure requirement; the Auto Scaling group must span multiple AZs, and the load balancer itself should be cross-zone enabled to distribute traffic across AZs.

407
MCQmedium

A trading dashboard stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured?

A.An EBS snapshot schedule
B.S3 Cross-Region Replication with versioning enabled
C.S3 lifecycle transition to Glacier Flexible Retrieval
D.A CloudFront distribution
AnswerB

Cross-Region Replication copies objects asynchronously to a bucket in another AWS Region, satisfying the disaster-recovery constraint. Versioning must be enabled on both source and destination because replication requires it to track object versions and propagate deletions or overwrites correctly.

Why this answer

S3 Cross-Region Replication (CRR) with versioning enabled automatically replicates objects to a destination bucket in a different AWS Region, providing a durable, low-latency disaster recovery copy. Versioning must be enabled on both source and destination buckets to track object changes and ensure consistency during replication. This meets the requirement for a cross-region copy without manual intervention.

Exam trap

The trap here is that candidates may confuse lifecycle transitions (which change storage class within the same region) with cross-region replication (which copies data to a different region), or assume EBS snapshots apply to S3 storage.

How to eliminate wrong answers

Option A is wrong because EBS snapshots are used for backing up EC2 block storage volumes, not for S3 objects, and they are region-specific unless manually copied. Option C is wrong because S3 lifecycle transition to Glacier Flexible Retrieval moves objects to a cold storage tier for cost savings, not to a different AWS Region for disaster recovery. Option D is wrong because CloudFront is a content delivery network that caches data at edge locations for low-latency access, not a mechanism for replicating data to another region for DR.

408
MCQeasy

An internal service is hosted behind an Application Load Balancer (ALB) with targets spread across two Availability Zones. If the targets in one Availability Zone become unhealthy, the service must continue serving traffic from the healthy AZ. What change most directly improves resilience at the load-balancing layer?

A.Turn off health checks and rely only on instance CPU utilization to route traffic.
B.Configure ALB listener rules to route all traffic to a single target group in one Availability Zone.
C.Configure target group health checks so the ALB stops sending traffic to unhealthy targets and continues routing to healthy targets in the other Availability Zone.
D.Store requests in an SQS queue before routing them to the ALB.
AnswerC

With target group health checks enabled and configured correctly, the ALB evaluates each target's health and stops routing requests to targets marked unhealthy. As long as healthy targets exist in the other AZ, the ALB preserves reachability.

Why this answer

Configuring target group health checks allows the ALB to automatically detect unhealthy targets and stop sending traffic to them, while continuing to route requests to healthy targets in the other Availability Zone. This directly improves resilience at the load-balancing layer by ensuring traffic is only forwarded to healthy instances, maintaining service availability even when an entire AZ fails.

Exam trap

The trap here is that candidates may think SQS decoupling (Option D) improves resilience at the load-balancing layer, but SQS operates at the application layer and does not affect how the ALB routes traffic to unhealthy targets.

How to eliminate wrong answers

Option A is wrong because turning off health checks removes the ALB's ability to detect unhealthy targets, which would cause traffic to be sent to failed instances, breaking resilience. Option B is wrong because routing all traffic to a single target group in one AZ creates a single point of failure and defeats the purpose of multi-AZ redundancy. Option D is wrong because storing requests in an SQS queue before routing to the ALB adds unnecessary latency and complexity, and does not address the immediate need for the ALB to stop sending traffic to unhealthy targets.

409
MCQmedium

A healthcare analytics company stores protected health information in an Amazon S3 bucket. An application running on Amazon EC2 instances in a private subnet must upload objects to the bucket using temporary credentials. The security team requires that the EC2 instances never store long-term AWS credentials on disk, and that access be limited to only the specific S3 bucket. Which solution meets these requirements?

A.Configure the S3 bucket policy to allow access from the EC2 instances' private IP addresses, and disable IAM authentication for the bucket.
B.Attach an IAM role to the EC2 instance profile with a policy granting s3:PutObject on the specific bucket, and let the AWS SDK retrieve temporary credentials automatically.
C.Generate a presigned URL for each upload using a role with broad S3 permissions, and distribute the URLs to the EC2 instances.
D.Create an IAM user with an access key, store the key in AWS Secrets Manager, and configure the application to retrieve it at runtime.
AnswerB

Attaching an IAM role to the instance profile allows the AWS SDK to obtain temporary credentials from the instance metadata service automatically. No long-term keys are stored on disk, and the attached policy can be scoped to grant only s3:PutObject on the specific bucket. This satisfies both the no-stored-credentials and least-privilege requirements with minimal operational overhead.

Why this answer

The requirement is temporary credentials without on-disk secrets, scoped to one bucket. An IAM role attached to the EC2 instance profile delivers rotating credentials through the instance metadata service, and the role's policy can be limited to the required S3 actions on the specific bucket. Long-term keys, presigned URLs, and IP-based policies do not satisfy the no-stored-credential and least-privilege conditions.

Exam trap

The trap here is assuming that storing long-term access keys in Secrets Manager converts them into temporary credentials.

410
Multi-Selecthard

A regional web application for a inventory service must fail over automatically to a secondary Region if the primary endpoint becomes unhealthy. Which two services or features are required? The design must avoid adding custom operational scripts.

Select 2 answers
A.Route 53 failover routing with health checks
B.S3 Transfer Acceleration
C.A deployed standby application stack in the secondary Region
D.AWS Organizations service control policies
AnswersA, C

Route 53 failover routing with health checks is the control-plane mechanism that implements active-passive regional failover. You configure a primary record with a health check that probes the primary endpoint's HTTP or TCP status, and a secondary record pointing to the standby region; when the health check fails, Route 53 stops returning the primary and starts resolving to the standby. The health check must be created separately and attached to the primary record set, and low TTLs or alias records are recommended so DNS clients quickly pick up the failed-over answer.

Why this answer

Route 53 failover routing with health checks (Option A) is required because it automatically evaluates the health of the primary endpoint and, upon detecting failure, updates DNS resolution to direct traffic to the secondary Region. This is the native AWS mechanism for DNS-based failover without custom scripts, relying on Route 53 health checkers to assess endpoint health via HTTP/HTTPS/TCP or calculated health checks.

Exam trap

The trap here is that candidates may think a single service like Route 53 alone can handle failover, but without a pre-deployed standby application stack in the secondary Region, there is no infrastructure to route traffic to, making both Route 53 failover routing and the standby stack required together.

411
MCQmedium

A media company runs a stateless transcoding fleet on Amazon EC2 instances spread across three Availability Zones behind a Network Load Balancer. The fleet must keep processing jobs even if an entire Availability Zone becomes unavailable, and the architect wants to minimize manual intervention. Which combination of actions should the architect take to meet these requirements?

A.Launch the instances as a Spot Fleet with a capacity-optimized allocation strategy and attach an Application Load Balancer in front of the fleet.
B.Create an Auto Scaling group that spans the three Availability Zones, enable ELB health checks, and configure the group to replace unhealthy instances automatically.
C.Configure the Network Load Balancer with cross-zone load balancing disabled and register all instances in a single target group.
D.Place all instances in a single Availability Zone and create an Amazon Route 53 latency routing policy pointing at the Network Load Balancer.
AnswerB

An Auto Scaling group spanning all three Availability Zones keeps capacity balanced across zones and, with ELB health checks enabled, automatically terminates and replaces instances that fail load balancer health checks. If one Availability Zone fails, the group launches replacement capacity in the remaining zones without operator involvement, which is exactly the resilience the scenario requires.

Why this answer

The requirement is automatic recovery from the loss of a whole Availability Zone. An Auto Scaling group distributed across three zones, combined with ELB health checks, continuously replaces unhealthy instances and rebalances capacity into surviving zones. That removes the need for manual intervention.

The other configurations either concentrate risk in one zone, rely on interruptible capacity, or alter load balancing behavior without providing replacement capacity.

Exam trap

The trap here is assuming that a load balancer by itself provides high availability, when in fact it only distributes traffic and cannot restore compute capacity after an Availability Zone fails.

412
MCQmedium

A static website uses an Amazon S3 bucket as the origin for an Amazon CloudFront distribution. The team accidentally configured the S3 bucket policy to allow s3:GetObject to Principal "*", so objects are accessible via direct S3 URLs. They want to ensure objects are retrievable only through CloudFront. What is the best corrective action?

A.Remove public access from the bucket and update the bucket policy to allow GetObject only from CloudFront using the distribution’s SourceArn (and use CloudFront origin access control or origin access identity).
B.Enable S3 static website hosting and disable CloudFront, because website hosting blocks direct object URL access.
C.Add a WAF rule that rate-limits requests to the S3 bucket domain to make direct access impractical.
D.Turn on S3 object versioning so that attackers cannot read previous objects.
AnswerA

Configure the S3 bucket to block all public access, then attach a bucket policy that grants s3:GetObject only to the CloudFront origin access control (OAC) identity, using the aws:SourceArn condition to restrict the principal to the exact CloudFront distribution ARN. This creates a hardened origin where the S3 bucket rejects any direct anonymous requests from the internet, while CloudFront's authenticated requests are permitted. Using OAC or OAI ensures the bucket owner has to intentionally authorize only that distribution, eliminating the common mistake of leaving objects publicly readable.

Why this answer

The S3 bucket policy currently allows s3:GetObject from any principal, making objects publicly accessible via direct S3 URLs. By removing public access and updating the policy to restrict GetObject to only requests that originate from the CloudFront distribution (using either Origin Access Control or Origin Access Identity), objects become retrievable exclusively through CloudFront, preventing direct S3 access.

Exam trap

The trap here is that candidates may think enabling S3 static website hosting or versioning solves the access control issue, but neither changes the bucket policy—only explicitly restricting the policy to CloudFront’s identity prevents direct S3 URL access.

How to eliminate wrong answers

Option B is wrong because enabling S3 static website hosting does not block direct object URL access; the S3 website endpoint is separate from the REST API endpoint, but the bucket policy still controls access, and objects remain accessible via direct S3 URLs unless the policy is restricted. Option C is wrong because a WAF rule applied to the S3 bucket domain is ineffective—WAF is a CloudFront feature and cannot be attached directly to an S3 bucket endpoint; rate-limiting would not prevent direct access, only reduce its frequency. Option D is wrong because enabling object versioning does not restrict access; it only preserves previous object versions, and without a restrictive bucket policy, all versions remain publicly accessible via direct S3 URLs.

413
MCQmedium

A public API for a customer analytics portal is deployed on API Gateway. Clients must authenticate with standards-based tokens issued by an external OpenID Connect provider. Which authorization mechanism should be used?

A.API keys only
B.JWT authorizer configured for the OpenID Connect issuer
C.IAM authorization for all internet users
D.A VPC endpoint policy
AnswerB

A JWT authorizer validates the token signature against the OpenID Connect issuer's published JWKS and checks claims such as audience and expiry. This satisfies the requirement that clients authenticate with standards-based tokens from an external OpenID Connect provider, which IAM authorization or API keys cannot validate.

Why this answer

The scenario requires standards-based token authentication from an external OpenID Connect (OIDC) provider. API Gateway's JWT authorizer natively validates JSON Web Tokens (JWTs) issued by OIDC providers by verifying the token's signature against the provider's JWKS endpoint, checking the `iss` and `aud` claims, and enforcing token expiration. This directly meets the requirement without needing custom Lambda authorizers or additional infrastructure.

Exam trap

The trap here is that candidates confuse API keys (which are static and not standards-based) with JWT tokens (which are cryptographically signed and verifiable), or assume IAM authorization can be used for external identities without understanding that IAM requires AWS credentials, not OIDC tokens.

How to eliminate wrong answers

Option A is wrong because API keys only provide simple identification and rate limiting, not authentication or authorization; they do not validate token signatures, claims, or issuer trust. Option C is wrong because IAM authorization is designed for AWS internal identities (IAM users/roles) and requires AWS Signature V4 signing, which is not compatible with external OIDC tokens or internet-based clients without custom signing logic. Option D is wrong because a VPC endpoint policy controls access to API Gateway via VPC endpoints, not authentication; it cannot validate OIDC tokens or handle client identity from the public internet.

414
MCQmedium

A company stores 500 TB of data in Amazon S3 Standard. The data is accessed frequently for the first 30 days after creation, then access drops to almost zero, but the data must be retained for 10 years for compliance. The company wants to minimize storage costs. Which solution is MOST cost-effective?

A.Enable S3 Versioning and use S3 Standard-Infrequent Access (S3 Standard-IA) after 30 days.
B.Use S3 One Zone-IA after 30 days to reduce costs.
C.Create an S3 Lifecycle policy to transition objects to S3 Glacier Deep Archive after 30 days.
D.Use S3 Intelligent-Tiering to automatically move data between access tiers.
AnswerC

S3 Glacier Deep Archive is the lowest-cost storage class for long-term retention, ideal for compliance data accessed rarely. Transitioning after 30 days aligns with the access pattern. Lifecycle policies automate the transition, eliminating manual intervention. This approach minimizes storage costs for the 10-year retention period, as Deep Archive costs significantly less than S3 Standard or other infrequent access tiers.

Why this answer

S3 Glacier Deep Archive offers the lowest storage cost for long-term retention and is designed for data accessed less than once per year. A lifecycle policy to transition after 30 days automates the process and ensures data is moved to the most economical tier. Other options either cost more or do not provide the necessary durability for compliance.

Exam trap

The trap here is assuming that S3 Intelligent-Tiering is always the most cost-effective for changing access patterns, but it incurs monitoring fees and is unnecessary when the access pattern is known and predictable.

415
MCQmedium

A company runs a two-tier web application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. The EC2 instances must access an Amazon Aurora MySQL DB cluster. A security engineer must ensure that only these EC2 instances can connect to the database, and that no credentials are stored on the instances. What should the security engineer do?

A.Create a VPC endpoint for Aurora and attach an IAM policy that allows only the EC2 instance role to use the endpoint.
B.Configure the DB cluster security group to allow inbound MySQL traffic from the EC2 instances' security group, and enable IAM database authentication on the cluster.
C.Store the Aurora master credentials in AWS Secrets Manager and grant the EC2 instance role permission to retrieve the secret at boot.
D.Place the DB cluster in a private subnet and attach an AWS WAF web ACL to the cluster endpoint to filter incoming connections.
AnswerB

Referencing the EC2 instances' security group as the source in the DB cluster security group allows only those instances to reach port 3306, and IAM database authentication lets the application generate a short-lived authentication token from its IAM role instead of storing a database password on the instance.

Why this answer

Restricting the Aurora security group to the EC2 instances' security group enforces network-level isolation so only those instances can open a MySQL session. Enabling IAM database authentication lets the application use its IAM role to obtain a temporary token, so no long-lived database password is stored on the instances, satisfying both the access and credential requirements.

Exam trap

The trap here is assuming that moving credentials into Secrets Manager also restricts which hosts can connect, when network access control and credential management are separate concerns.

416
MCQeasy

An event consumer sometimes processes the same SQS message more than once due to timeouts and retries. The consumer must ensure the payment is not charged twice. What design choice best addresses this requirement?

A.Assume messages are processed exactly once because SQS uses durable storage.
B.Make the payment operation idempotent by using an idempotency key and skipping side effects when the key indicates the payment already succeeded.
C.Increase the consumer visibility timeout to several days so messages are not redelivered.
D.Delete the message immediately even if processing fails validation.
AnswerB

Idempotency ensures that repeated processing attempts produce the same result. The consumer should use a stable idempotency key (for example, a business transaction ID) and record completion in durable storage. If the key already indicates the payment succeeded, the consumer skips charging again.

Why this answer

Making the payment operation idempotent using an idempotency key ensures that even if the same SQS message is processed multiple times due to timeouts and retries, the payment will only be charged once. The consumer checks the idempotency key before executing the payment; if the key indicates the payment already succeeded, the consumer skips the side effect. This pattern directly addresses the requirement of not charging twice without relying on SQS's at-least-once delivery guarantee.

Exam trap

The trap here is that candidates assume SQS provides exactly-once delivery or that increasing the visibility timeout is a reliable solution, but the exam tests understanding that SQS is at-least-once and that idempotency is the correct architectural pattern to handle duplicates.

How to eliminate wrong answers

Option A is wrong because SQS guarantees at-least-once delivery, not exactly-once processing; messages can be duplicated due to network issues or consumer timeouts, so assuming exactly-once processing is incorrect. Option C is wrong because increasing the visibility timeout to several days does not prevent redelivery; it only delays it, and if the consumer crashes or fails to delete the message, it will still be redelivered after the timeout expires. Option D is wrong because deleting a message immediately even if processing fails validation means the message is lost permanently, preventing any retry or dead-letter queue handling, which can lead to data loss or incomplete processing.

417
MCQmedium

A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The architecture review board prefers a managed AWS-native control.

A.Lambda reserved concurrency set to zero
B.A Lambda dead-letter queue or failure destination
C.A larger deployment package
D.CloudFront error pages
AnswerB

Configuring a Lambda dead-letter queue (DLQ) or an asynchronous failure destination is the correct solution because Lambda will automatically send events that have exhausted built-in retries to a specified SQS queue, SNS topic, or other destination. This preserves the failed event payload and metadata, enabling you to inspect, replay, or alert on the failure. For asynchronous invocations, this is the standard pattern for building resilient, observable serverless workflows.

Why this answer

Lambda dead-letter queues (DLQs) or failure destinations are the correct AWS-native mechanism to retain failed events after retries are exhausted. When a Lambda function fails to process an event (e.g., due to an unreliable third-party API), the function can be configured to send the failed event payload to an SQS queue or SNS topic for later investigation. This ensures no data loss and aligns with the requirement for a managed, AWS-native solution.

Exam trap

The trap here is that candidates often confuse Lambda DLQs with SQS DLQs or assume that increasing retries (via reserved concurrency or package size) solves the retention problem, but the key is the explicit configuration to capture events after retries are exhausted.

How to eliminate wrong answers

Option A is wrong because setting Lambda reserved concurrency to zero would prevent the function from executing at all, not retain failed events. Option C is wrong because a larger deployment package has no impact on error handling or event retention; it only affects cold start times and deployment size. Option D is wrong because CloudFront error pages are for HTTP-level errors in front of web applications, not for Lambda function invocation failures or event retention.

418
MCQeasy

A company’s private workload in a VPC uploads objects to an S3 bucket. Security requires that S3 requests are allowed only when they traverse a specific S3 Gateway VPC Endpoint (vpce-0abc123example). Which change best enforces this restriction at the S3 bucket level?

A.Add an S3 bucket policy Deny statement for s3:PutObject when aws:sourceVpce is not equal to vpce-0abc123example.
B.Add an S3 bucket policy Deny statement that blocks requests unless the principal uses MFA.
C.Enable Block Public Access and remove the public bucket policy statement.
D.Attach an IAM policy to the workload role that allows s3:PutObject only to the bucket ARN.
AnswerA

A bucket policy can use the request context key aws:sourceVpce to distinguish requests that came through a particular VPC endpoint. Using a Deny with a condition such as StringNotEquals on aws:sourceVpce blocks PutObject unless the request reached S3 via that specific Gateway Endpoint. Requests that arrive by other network paths will not match the required endpoint ID and will be denied.

Why this answer

It uses an S3 bucket policy with a Deny statement that explicitly denies any s3:PutObject request unless the request originates from the specified VPC Endpoint (vpce-0abc123example). The aws:sourceVpce condition key evaluates the VPC endpoint ID from which the request is made, ensuring that only traffic through that specific Gateway VPC Endpoint is allowed. This enforces the security requirement at the bucket level, overriding any other policies that might allow access from other sources.

Exam trap

The trap here is that candidates often confuse IAM policies (which control who can act) with bucket policies (which control how and from where access is allowed), leading them to choose an IAM-based solution (Option D) that does not enforce the network-level restriction required by the scenario.

How to eliminate wrong answers

Option B is wrong because requiring MFA does not restrict requests to a specific VPC Endpoint; it only adds an authentication factor, which does not enforce the network-level restriction. Option C is wrong because Block Public Access and removing public policies prevent public access but do not restrict requests to a specific VPC Endpoint; private traffic from other sources (e.g., the internet via a NAT gateway) would still be allowed. Option D is wrong because an IAM policy attached to the workload role controls what the role can do but does not restrict the network path; the workload could still send requests from any network interface, not just the specified VPC Endpoint.

419
MCQmedium

A media company runs a transcoding fleet on Amazon EC2 instances behind an Application Load Balancer in a single Availability Zone. The business requires the workload to survive the loss of that Availability Zone with no manual intervention and minimal downtime. The instances store intermediate files on instance store volumes and the fleet is managed by an Auto Scaling group. Which change should a solutions architect make to meet the requirement?

A.Increase the desired capacity of the Auto Scaling group so that more instances run in the existing Availability Zone.
B.Convert the instances to Spot Instances and configure a placement group to keep them close together.
C.Add a second Availability Zone to the Auto Scaling group and configure the Application Load Balancer with subnets in both Availability Zones.
D.Enable detailed CloudWatch monitoring and create an alarm that notifies operators when an instance becomes unhealthy.
AnswerC

Distributing the Auto Scaling group across two Availability Zones and attaching the load balancer to subnets in both zones lets the fleet continue serving traffic if one zone fails. The load balancer health checks remove unhealthy instances automatically, and the group replaces them, so the workload survives the loss without manual action.

Why this answer

Resilience to an Availability Zone failure requires spreading compute across multiple zones and giving the load balancer subnets in each of those zones. The Auto Scaling group then replaces failed instances in the surviving zone while the load balancer routes only to healthy targets, delivering continuous service without human action.

Exam trap

The trap here is assuming that adding more instances inside one Availability Zone provides high availability, when zone-level isolation is what actually protects the workload.

420
MCQmedium

A company runs an EC2 Auto Scaling group behind an internet-facing Application Load Balancer. The security team must ensure that the instances accept HTTP traffic only from the ALB and never directly from the internet, while the ALB itself must accept traffic only from a specific corporate CIDR range. Which combination of security group configurations should a solutions architect implement?

A.Create an inbound rule on the instance security group allowing TCP 80 from 0.0.0.0/0, and an inbound rule on the ALB security group allowing TCP 80 from the corporate CIDR range.
B.Create an inbound rule on the instance security group allowing TCP 80 from the ALB's security group, and an inbound rule on the ALB security group allowing TCP 80 from the corporate CIDR range.
C.Create an inbound rule on the instance security group allowing TCP 80 from the corporate CIDR range, and an inbound rule on the ALB security group allowing TCP 80 from the ALB's security group.
D.Create an inbound rule on the instance security group allowing TCP 80 from the corporate CIDR range, and an inbound rule on the ALB security group allowing TCP 80 from 0.0.0.0/0.
AnswerB

Referencing the ALB's security group as the source means only traffic that the ALB forwards can reach the instances, regardless of the instances' public IPs. Restricting the ALB security group to the corporate CIDR range enforces the second requirement. This is the standard two-tier security group pattern and satisfies both conditions without extra infrastructure.

Why this answer

The secure pattern is to make the instances trust only the load balancer by referencing the ALB's security group as the source, and to make the ALB trust only the intended clients by scoping its inbound rule to the corporate CIDR range. This creates a chained trust model where the instances never accept direct internet connections and the ALB is not exposed beyond the approved network range.

Exam trap

The trap here is assuming the instances can match on the original client CIDR, when the ALB rewrites the source address so only a security group reference reliably identifies the load balancer.

421
Multi-Selecthard

A partner integration sends a custom binary TCP protocol to a service running on EC2 instances in private subnets. The partners require static endpoint IPs for allowlisting, and the application must see the original client source IP for rate limiting. Which two changes best fit the protocol and network requirements? Select two.

Select 2 answers
A.Replace the Application Load Balancer with a Network Load Balancer.
B.Use a TCP listener on the load balancer instead of an HTTP or HTTPS listener.
C.Put the service behind API Gateway REST API and use Lambda integration.
D.Use CloudFront to cache the binary packets at edge locations.
E.Terminate the traffic with an Amazon RDS proxy to stabilize the connections.
AnswersA, B

A Network Load Balancer is the right choice for TCP traffic and low-latency forwarding at layer 4. It also supports static IP behavior that is important for partner allowlisting. This directly matches the custom binary protocol and source-IP requirement.

Why this answer

A Network Load Balancer (NLB) is required because it supports TCP traffic natively at Layer 4, which is necessary for a custom binary TCP protocol that cannot be interpreted by an Application Load Balancer (ALB) at Layer 7. Additionally, an NLB preserves the original client source IP address by default when used with targets in private subnets, meeting the requirement for rate limiting based on the client IP. Static IP addresses can be assigned to the NLB via Elastic IPs, satisfying the partner's need for static endpoint IPs for allowlisting.

Exam trap

The trap here is that candidates often assume an Application Load Balancer can handle any TCP traffic because it supports TCP listeners, but ALB only supports HTTP/HTTPS at Layer 7 and cannot process custom binary protocols, while NLB is the correct choice for non-HTTP TCP traffic with static IP and client IP preservation requirements.

422
MCQmedium

A media company runs a video transcoding pipeline on Amazon EC2 instances in a single Availability Zone. The pipeline writes intermediate files to an Amazon EBS volume attached to each instance. The company needs the pipeline to survive the failure of any single Availability Zone and to recover automatically with minimal data loss. Which change should a solutions architect make?

A.Create an Auto Scaling group that spans multiple Availability Zones and configure each instance to store intermediate files on an instance store volume.
B.Create an Auto Scaling group that spans multiple Availability Zones and use Amazon EBS Multi-Attach volumes shared across all instances.
C.Create an Auto Scaling group in one Availability Zone and take scheduled EBS snapshots every hour to another Availability Zone.
D.Create an Auto Scaling group that spans multiple Availability Zones and store intermediate files in an Amazon S3 bucket.
AnswerD

Amazon S3 is a regional service with high durability, so intermediate files survive an Availability Zone failure. An Auto Scaling group spanning multiple Availability Zones replaces failed instances automatically, and new instances can read the intermediate files from the S3 bucket. This meets both the resilience and minimal data loss requirements.

Why this answer

Storing intermediate files in Amazon S3 decouples the pipeline from any single Availability Zone because S3 is a regional, highly durable service. An Auto Scaling group spanning multiple Availability Zones automatically replaces instances when a zone fails, and replacement instances retrieve the intermediate files from S3, minimizing data loss and manual recovery effort.

Exam trap

The trap here is assuming that an Auto Scaling group spanning multiple Availability Zones alone provides resilience, when the storage layer must also be decoupled from a single zone.

423
MCQeasy

A company runs a steady-state web application on a fixed number of Amazon EC2 instances that have been running continuously for over a year. The workload is predictable and will remain in production for at least three more years. Management wants to reduce compute cost without changing the architecture. Which purchasing option should a solutions architect recommend?

A.Purchase a 3-year Compute Savings Plan with a partial upfront payment for the steady EC2 usage.
B.Use On-Demand Instances and enable detailed monitoring to improve utilization visibility.
C.Move the application to Dedicated Instances to obtain volume discounts on the hourly rate.
D.Convert the instances to Spot Instances to take advantage of unused EC2 capacity.
AnswerA

A Compute Savings Plan applies to EC2 usage regardless of instance family, size, tenancy, or Region, and a 3-year term with partial upfront payment yields a significant discount over On-Demand. Because the workload is predictable and will run for at least three more years, this commitment matches the usage and reduces compute cost without architectural changes.

Why this answer

For predictable, long-running EC2 usage, a Compute Savings Plan provides a lower effective hourly rate in exchange for a term commitment. It is flexible across instance families and Regions, so it reduces cost without requiring the application to change. Spot risks interruption, On-Demand forgoes discounts, and Dedicated Instances raise cost rather than lower it.

Exam trap

The trap here is choosing Spot because it has the deepest discount, while ignoring that a steady production web application cannot tolerate the interruptions Spot permits.

424
MCQmedium

A telemetry pipeline uses an Application Load Balancer in one Region. Global users need lower network latency to the application without caching dynamic responses. What should be considered?

A.AWS Global Accelerator
B.S3 Cross-Region Replication
C.CloudFront only with long TTLs
D.AWS Backup cross-Region copy
AnswerA

AWS Global Accelerator uses static anycast IP addresses at AWS edge locations to accept TCP/UDP traffic and then forwards it over the AWS global backbone network to the optimal regional endpoint (e.g., an Application Load Balancer) rather than traversing the public internet. For a telemetry pipeline with many dynamic requests, this reduces latency, jitter, and packet loss while providing fast failover via health checks—benefits that come from network path optimization, not caching.

Why this answer

AWS Global Accelerator uses the Anycast IP address concept to route traffic through the AWS global network to the optimal endpoint, reducing latency and jitter for global users. It does not cache content, making it ideal for dynamic responses that cannot be cached, and it integrates directly with an Application Load Balancer in a single Region.

Exam trap

The trap here is that candidates often confuse CloudFront (a CDN with caching) with Global Accelerator (a non-caching network accelerator), assuming any edge service must cache content, but Global Accelerator is designed specifically for dynamic and uncacheable traffic.

How to eliminate wrong answers

Option B is wrong because S3 Cross-Region Replication is a storage feature for replicating objects across S3 buckets in different Regions; it does not reduce network latency for application traffic or handle dynamic HTTP responses. Option C is wrong because CloudFront with long TTLs caches responses at edge locations, which is unsuitable for dynamic content that must not be cached; long TTLs would serve stale data. Option D is wrong because AWS Backup cross-Region copy is a disaster recovery feature for backing up resources to another Region; it does not improve real-time network latency for users accessing the application.

425
MCQmedium

A SaaS company uses an S3 bucket for database backups created daily. Backups are rarely restored; the company’s documented RTO is 24 hours, and the compliance policy requires backups be kept for 90 days. The team currently stores all backups in S3 Standard, which is costly. Which single lifecycle policy change is most cost-optimized while still meeting the 24-hour RTO and 90-day retention?

A.Add a lifecycle rule to transition backups older than 1 day to S3 Glacier Flexible Retrieval, and keep them until day 90.
B.Add a lifecycle rule to transition backups older than 1 day to S3 Glacier Instant Retrieval, and keep them until day 90.
C.Add a lifecycle rule to transition backups older than 1 day to S3 Glacier Deep Archive, and keep them until day 90 with no restore configuration.
D.Add a lifecycle rule to transition backups older than 1 day to S3 One Zone-IA, and delete them after 7 days.
AnswerA

This option correctly leverages S3 Lifecycle rules to transition older, less frequently accessed backups to S3 Glacier Flexible Retrieval. This storage class provides significant cost savings compared to S3 Standard or S3-IA, while still supporting retrieval times measured in hours, which comfortably meets a 24-hour Recovery Time Objective (RTO). Maintaining retention until day 90 also satisfies the long-term data retention requirement efficiently.

Why this answer

S3 Glacier Flexible Retrieval provides retrieval times from minutes to hours, which meets the 24-hour RTO, and offers significant cost savings over S3 Standard for data that is rarely accessed. Transitioning backups older than 1 day to this storage class reduces costs while retaining them for the required 90-day compliance period.

Exam trap

The trap here is that candidates may choose S3 Glacier Deep Archive for maximum cost savings without verifying that its retrieval time (12–48 hours) can exceed the 24-hour RTO, or they may overlook that S3 Glacier Instant Retrieval is not the most cost-effective option for data that is restored only rarely.

How to eliminate wrong answers

Option B is wrong because S3 Glacier Instant Retrieval is designed for data accessed once a quarter with millisecond retrieval, but it is more expensive than S3 Glacier Flexible Retrieval and not the most cost-optimized choice for backups restored only rarely within a 24-hour RTO. Option C is wrong because S3 Glacier Deep Archive has a retrieval time of 12–48 hours, which may exceed the 24-hour RTO, and the option lacks a restore configuration, making it non-compliant with the RTO requirement. Option D is wrong because S3 One Zone-IA does not provide the durability or availability needed for critical backups, and deleting backups after 7 days violates the 90-day retention policy.

426
MCQmedium

A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured?

A.Lambda reserved concurrency set to zero
B.A Lambda dead-letter queue or failure destination
C.A larger deployment package
D.CloudFront error pages
AnswerB

Configuring a dead-letter queue (an SQS queue or SNS topic) or an asynchronous failure destination causes Lambda to route the event to that target after the configured retry attempts are exhausted. The original event payload is preserved in the DLQ or delivered to a destination such as SQS, SNS, EventBridge, or another Lambda function, allowing operators to inspect and later reprocess the failed inventory events. This is the correct mechanism for capturing failed asynchronous invocations for analysis.

Why this answer

Lambda dead-letter queues (DLQs) or failure destinations are the correct mechanism to retain failed events after all retries are exhausted. When a Lambda function fails to process an event (e.g., from an asynchronous invocation), the service automatically retries twice. If those retries fail, the event can be sent to an SQS queue or SNS topic (DLQ) or to a specified destination (failure destination) for later investigation.

This ensures no data loss and provides a durable storage for post-mortem analysis.

Exam trap

The trap here is that candidates may confuse DLQs with retry mechanisms or think that increasing function resources (like memory or package size) will prevent failures, when in fact DLQs are the only way to durably capture events after retries are exhausted.

How to eliminate wrong answers

Option A is wrong because setting reserved concurrency to zero would prevent the Lambda function from executing at all, not retain failed events. Option C is wrong because a larger deployment package does not affect error handling or event retention; it only increases cold start latency and storage overhead. Option D is wrong because CloudFront error pages are for HTTP-level errors from a web distribution, not for capturing asynchronous Lambda invocation failures.

427
MCQmedium

Based on the exhibit, the application team wants the database to keep the same connection endpoint during failover and to reconnect automatically after the primary instance becomes unavailable. Which change best meets the requirement?

A.Keep the IP address and increase the JDBC connection timeout so the application waits longer during failover.
B.Replace the IP address with the RDS DNS endpoint and add client retry logic that re-resolves DNS after connection loss.
C.Create an additional read replica and point the application to it so failover is faster.
D.Place a Network Load Balancer in front of the database and use the load balancer target IP to avoid DNS changes.
AnswerB

RDS Multi-AZ failover preserves the database endpoint name, not the underlying IP address. When the standby is promoted, AWS updates the DNS record to point to the new primary. Using the RDS endpoint allows the application to follow that change, and retry logic helps the client recover from the short disconnect that occurs during failover.

Why this answer

Using the RDS DNS endpoint ensures that the application connects to the current primary instance, even after a failover. When the primary becomes unavailable, RDS promotes a standby (or read replica) to a new primary and updates the DNS record to point to the new instance's IP. By adding client retry logic that re-resolves DNS after a connection loss, the application automatically picks up the new IP and reconnects without manual intervention, meeting both requirements of a stable endpoint and automatic reconnection.

Exam trap

The trap here is that candidates assume a static IP or a load balancer can provide a stable endpoint, but AWS RDS does not support static IPs for Multi-AZ failover, and NLB cannot front RDS instances—the only reliable way is to use the RDS DNS endpoint with retry logic that re-resolves DNS after a connection loss.

How to eliminate wrong answers

Option A is wrong because keeping the IP address is unreliable—after a failover, the new primary instance will have a different IP address, so the application would connect to a stale IP and fail. Increasing the JDBC connection timeout only delays the failure; it does not resolve the underlying IP mismatch. Option C is wrong because creating an additional read replica does not change the connection endpoint for the primary; the application still connects to the original primary endpoint, which becomes unavailable during failover.

Read replicas are for read scaling, not for providing a failover endpoint. Option D is wrong because placing a Network Load Balancer in front of an RDS database is not a supported architecture—RDS does not integrate with NLB for database traffic, and the load balancer target IP would still change after failover, requiring DNS re-resolution anyway, making the solution unnecessarily complex and non-compliant with AWS best practices.

428
MCQmedium

A distributed system needs extremely low network latency between a set of EC2 instances running the same workload. The team wants the instances to be placed as close together as AWS allows to reduce round-trip time. Which placement strategy should the architect use?

A.Use a Cluster placement group for the instances that must communicate frequently over low latency.
B.Use a Spread placement group across multiple Availability Zones to maximize fault tolerance.
C.Use the default placement strategy without specifying a placement group.
D.Use a placement group of type Partition to ensure independent failure of each instance.
AnswerA

Cluster placement groups are designed to place instances close together within a single Availability Zone to minimize network latency. They are the right choice when nodes require high intercommunication performance, such as distributed processing or tightly coupled systems. The scenario’s goal of minimizing round-trip time aligns with the Cluster placement group behavior. It’s also an EC2-native placement option focused on performance.

Why this answer

A Cluster placement group is the correct choice because it places instances in a single Availability Zone within the same rack or logical cluster, providing the lowest possible network latency and maximum throughput (up to 10 Gbps for single-flow traffic) between instances. This is ideal for tightly coupled, latency-sensitive workloads like HPC or real-time distributed systems.

Exam trap

The trap here is that candidates often confuse the purpose of placement groups: Cluster is for low latency and high throughput, Spread is for fault tolerance across hardware, and Partition is for large distributed systems needing failure isolation, but only Cluster guarantees physical proximity.

How to eliminate wrong answers

Option B is wrong because a Spread placement group spreads instances across distinct hardware racks or Availability Zones, which increases latency and is designed for fault tolerance, not low latency. Option C is wrong because the default placement strategy does not guarantee proximity; instances may be placed on different racks or AZs, leading to higher latency. Option D is wrong because a Partition placement group spreads instances across multiple partitions (each with separate racks) to isolate failures, but does not minimize latency between instances within the same partition.

429
MCQmedium

A content publishing system uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The architecture review board prefers a managed AWS-native control.

A.Lambda reserved concurrency set to zero
B.A larger deployment package
C.CloudFront error pages
D.A Lambda dead-letter queue or failure destination
AnswerD

A dead-letter queue (SQS/SNS) or a failure destination is the correct mechanism because Lambda, after exhausting its default two retries, can route the original event payload to a configured target. A DLQ preserves the raw event and allows a separate process to consume, inspect, and reprocess it, while a failure destination offers richer metadata, such as the request ID and response context, and can send to SQS, SNS, Lambda, or EventBridge. This gives a durable record of failed async invocations and decouples error handling from the main processing function.

Why this answer

Lambda dead-letter queues (DLQs) or failure destinations are the managed AWS-native way to capture events that have exhausted all retry attempts from an asynchronous invocation. When the Lambda function fails after the configured number of retries (default 3), the event is automatically sent to an SQS queue or SNS topic (DLQ) or to a specified destination (e.g., SQS, SNS, EventBridge) for later investigation and reprocessing.

Exam trap

The trap here is that candidates may confuse Lambda's synchronous invocation retry behavior (which is controlled by the caller) with asynchronous invocation retries (which are managed by Lambda itself and require a DLQ or failure destination for post-retry capture).

How to eliminate wrong answers

Option A is wrong because setting reserved concurrency to zero would prevent the Lambda function from executing at all, not handle failed events after retries. Option B is wrong because a larger deployment package does not affect retry or failure handling; it only increases the function's code size and cold start latency. Option C is wrong because CloudFront error pages are for HTTP-level errors from a web distribution, not for capturing failed asynchronous Lambda invocations from a third-party API call.

430
MCQhard

A company runs a microservices application on Amazon ECS with AWS Fargate. The services communicate over HTTP/2 and gRPC. The architect needs to implement service-to-service communication that provides high throughput, low latency, and mutual TLS encryption. The solution must also support traffic splitting for canary deployments. Which approach should the architect take?

A.Use an Application Load Balancer with gRPC support and configure multiple target groups with weighted routing.
B.Use Amazon API Gateway with private integration to the ECS services, and enable mutual TLS on the API Gateway.
C.Deploy an AWS App Mesh virtual service with Envoy proxies sidecar containers, and configure routes with weights for canary deployments.
D.Configure AWS Cloud Map for service discovery and use security groups to enforce encryption between tasks.
AnswerC

AWS App Mesh is a service mesh that provides consistent networking for microservices. It uses Envoy proxies to handle service-to-service communication, supporting HTTP/2 and gRPC. It can enforce mutual TLS for encryption and provides traffic routing capabilities, including weighted targets for canary deployments. This meets all requirements: high throughput, low latency, mTLS, and traffic splitting.

Why this answer

AWS App Mesh is a service mesh that uses Envoy proxies to manage service-to-service communication. It supports HTTP/2 and gRPC, provides mutual TLS for encryption, and allows weighted routing for canary deployments. This makes it the best fit for the requirements of high throughput, low latency, mTLS, and traffic splitting in an ECS on Fargate environment.

Exam trap

The trap here is thinking that an Application Load Balancer or API Gateway can provide mutual TLS for service-to-service communication, when they are primarily for north-south traffic and do not offer the same mesh capabilities.

431
MCQmedium

A healthcare company stores patient imaging studies in an Amazon S3 bucket encrypted with SSE-KMS using a customer managed key. A security audit reveals that a former employee's IAM user still has s3:GetObject permissions on the bucket. The company wants to ensure the former employee can no longer decrypt any objects, even if they somehow regain S3 access, without affecting other users or applications. What should a security engineer do?

A.Disable the customer managed KMS key used for the bucket encryption.
B.Add a bucket policy that denies s3:GetObject to the former employee's IAM user ARN.
C.Enable S3 Block Public Access on the bucket and rotate the IAM user's access keys.
D.Update the KMS key policy to explicitly deny the former employee's IAM user the kms:Decrypt action.
AnswerD

SSE-KMS requires the caller to have kms:Decrypt permission on the customer managed key in addition to s3:GetObject. Adding an explicit deny for the former employee in the key policy guarantees they cannot decrypt objects even if IAM or bucket policies later grant S3 access, because an explicit deny in the key policy overrides any allow. This achieves targeted revocation without impacting other principals.

Why this answer

With SSE-KMS, decryption requires both S3 read permission and kms:Decrypt on the customer managed key. Adding an explicit deny for the former employee in the KMS key policy ensures they cannot decrypt objects even if S3 access is accidentally restored, while leaving other principals unaffected. Bucket-level or access key actions do not remove the cryptographic capability, and disabling the key would break access for everyone.

Exam trap

The trap here is assuming that removing S3 permissions alone is sufficient, when SSE-KMS decryption also requires kms:Decrypt authorization on the key.

432
MCQeasy

A team stores application logs in Amazon S3. They need access to the logs only occasionally for troubleshooting (infrequent access), and they want to reduce storage cost automatically over time without manually moving objects. What should they implement?

A.An S3 lifecycle policy that transitions objects to a lower-cost storage class after a set number of days
B.An S3 lifecycle policy that deletes objects after 1 day to eliminate storage costs
C.An S3 lifecycle policy that keeps all objects in S3 Standard and only applies compression at read time
D.A policy that changes bucket encryption from SSE-S3 to SSE-KMS to reduce storage cost
AnswerA

S3 lifecycle policies can automatically transition objects based on age to storage classes priced for infrequent access (for example, Standard-IA or Glacier-based classes). This preserves the data for later troubleshooting while lowering storage cost as objects become older.

Why this answer

An S3 lifecycle policy can automatically transition objects from S3 Standard to lower-cost storage classes (e.g., S3 Standard-IA, S3 One Zone-IA, or S3 Glacier Instant Retrieval) after a specified number of days. This meets the requirement of reducing storage costs over time for infrequently accessed logs without manual intervention, as the policy automates the movement based on object age.

Exam trap

The trap here is that candidates may confuse lifecycle policies with deletion policies, thinking that deleting objects after a short period (Option B) is a valid cost-saving strategy, but the question explicitly requires retaining logs for occasional troubleshooting, so deletion is not appropriate.

Why the other options are wrong

B

Deleting logs after 1 day prevents troubleshooting access for incidents that may occur later, and the requirement is to reduce cost automatically over time, not eliminate data immediately.

C

S3 Lifecycle policies cannot apply compression at read time; they only transition or expire objects. Compression at read time is not a storage cost reduction feature and would not automatically reduce costs over time.

D

Changing encryption from SSE-S3 to SSE-KMS does not reduce storage cost; in fact, SSE-KMS incurs additional KMS key usage charges, increasing cost.

When would these options actually be correct?

B

An S3 lifecycle policy that deletes objects after a set number of days would be correct if the requirement is to automatically remove logs after a retention period (e.g., compliance mandates deletion after 90 days) and no longer-term access is needed.

C

An S3 Lifecycle policy that keeps all objects in S3 Standard and applies compression at read time would be correct if the question asked for a solution to reduce data transfer costs or improve read performance for frequently accessed logs, while maintaining immediate access without storage class transitions.

D

A question where the requirement is to enforce customer-managed encryption keys for compliance or to control key rotation, and cost is not the primary concern.

Why candidates pick the wrong answer

B

Candidates may think deleting old logs is the simplest way to reduce cost, overlooking the need for occasional troubleshooting access and the availability of lower-cost storage classes.

C

Candidates may think compression reduces storage costs, but S3 does not support server-side compression at read time, and the option incorrectly implies that compression can be applied automatically via a lifecycle policy.

D

Candidates may mistakenly believe that using a more advanced encryption method like KMS reduces storage costs, or they confuse encryption cost with storage cost.

433
MCQmedium

A internal reporting portal serves infrequently accessed user documents that must be available immediately when requested. Which S3 storage class is likely the best cost fit?

A.Instance store volumes
B.S3 Glacier Deep Archive
C.S3 Standard for all objects
D.S3 Standard-IA or S3 One Zone-IA depending on resilience requirements
AnswerD

S3 Standard-IA is the right baseline for infrequent access because it charges less for storage than S3 Standard while still providing millisecond first-byte latency and multi-AZ durability. If the reporting data can be regenerated or an AZ failure is an acceptable risk, S3 One Zone-IA costs even less by storing copies in a single availability zone, but it sacrifices resilience to AZ destruction. The choice between the two depends on the portal's recovery point objective and whether losing the current copy would be disruptive to the business.

Why this answer

S3 Standard-IA or S3 One Zone-IA is the best cost fit because the data is infrequently accessed but must be available immediately when requested. These storage classes offer low-latency retrieval (milliseconds) at a lower storage cost than S3 Standard, with the trade-off of a retrieval fee. The choice between Standard-IA and One Zone-IA depends on whether the application requires resilience against Availability Zone failures.

Exam trap

The trap here is that candidates may choose S3 Standard for all objects because they assume 'immediately available' requires the highest performance tier, overlooking that Standard-IA and One Zone-IA offer identical retrieval latency at a lower storage cost for infrequently accessed data.

How to eliminate wrong answers

Option A is wrong because instance store volumes are ephemeral block storage attached to EC2 instances, not an S3 storage class, and data is lost if the instance is stopped or terminated. Option B is wrong because S3 Glacier Deep Archive has retrieval times of 12–48 hours, which does not meet the 'immediately available' requirement. Option C is wrong because S3 Standard is designed for frequently accessed data and would incur higher storage costs for infrequently accessed objects, making it less cost-optimal than Standard-IA or One Zone-IA.

434
MCQmedium

A backup process restores a 2 TB production database from an EBS snapshot onto a new volume. During the first hours after restore, the application sees slow reads whenever previously unused blocks are accessed. What is the best way to avoid this performance issue in future restores?

A.Increase the volume size to give the database more free space.
B.Enable Fast Snapshot Restore on the snapshots used for recovery.
C.Move the database files to Amazon EFS after the restore completes.
D.Use magnetic standard volumes because they avoid snapshot hydration delays.
AnswerB

Fast Snapshot Restore removes the initial performance penalty that occurs when a restored EBS volume reads blocks that have not yet been hydrated. By pre-warming the snapshot data in the target AZ, it helps ensure consistent read performance immediately after restore. This is especially valuable for databases and other workloads that must recover quickly without waiting for the background hydration process.

Why this answer

When an EBS volume is restored from a snapshot, it is lazily loaded from Amazon S3 in the background. Accessing data blocks that have not yet been loaded triggers a read penalty because the volume must fetch them from S3 before serving the I/O. Enabling Fast Snapshot Restore (FSR) pre-warms the snapshot data so that restored volumes have full performance immediately, eliminating the slow reads on first access.

Exam trap

The trap here is that candidates may think increasing volume size or switching to a different storage class will fix the lazy hydration delay, but only Fast Snapshot Restore directly addresses the root cause by pre-initializing the data blocks.

Why the other options are wrong

A

Increasing volume size does not address the 'first touch' latency caused by lazy loading of data from snapshot to S3; it only provides more storage capacity.

C

Moving database files to Amazon EFS after restore does not address the slow reads caused by lazy loading of data from EBS snapshots (snapshot hydration). EFS is a network file system with its own performance characteristics and does not eliminate the need to initialize EBS blocks.

D

Magnetic standard volumes (st1/sc1) also suffer from snapshot hydration delays and have lower baseline performance than gp2/gp3, making them unsuitable for avoiding slow reads on previously unused blocks.

When would these options actually be correct?

A

A scenario where the database is running out of storage space and experiencing performance degradation due to insufficient IOPS or throughput, and increasing volume size would also increase the volume's baseline performance.

C

A scenario where the application requires shared access to the database files across multiple EC2 instances, or where the database needs to be accessed from different Availability Zones for high availability, and the performance impact of snapshot hydration is acceptable or mitigated by other means.

D

A question asks for the most cost-effective storage for a large, sequential-access data warehouse that is rarely accessed and can tolerate lower IOPS. Magnetic volumes would be correct due to their low cost per GB.

Why candidates pick the wrong answer

A

Candidates may think larger volumes inherently perform better or that more free space reduces fragmentation, but the issue is specifically about snapshot hydration delays, not capacity.

C

Candidates may think that using a different storage service like EFS, which is fully managed and scalable, could bypass the EBS snapshot hydration issue, not realizing that the problem is specific to EBS volumes restored from snapshots.

D

Candidates may think that older, simpler technology (magnetic) avoids the 'hydration' issue because they misunderstand that all EBS snapshots lazily restore blocks, regardless of volume type.

435
MCQeasy

Your global users access static images stored in S3. Origin bandwidth costs are higher than expected because CloudFront is not caching effectively. What change most directly reduces origin fetches (and typically lowers data transfer costs) without changing application logic?

A.Configure CloudFront caching by setting appropriate cache-control headers and/or CloudFront cache policy/TTL values for the static objects
B.Disable CloudFront caching so every request goes back to S3 for the latest image
C.Route users directly to the S3 website endpoint to bypass CloudFront
D.Turn on a NAT Gateway for the CloudFront origin to reduce bandwidth charges
AnswerA

CloudFront reduces origin fetches when responses are cacheable and allowed to remain in the edge cache for a meaningful duration. Ensuring the objects include correct cache-control headers (or configuring CloudFront cache policy TTLs) increases cache hit rate, so fewer requests require fetching from S3 origin. This directly reduces origin bandwidth and related data transfer costs.

Why this answer

The high origin bandwidth costs are caused by CloudFront not caching effectively, meaning too many requests reach the S3 origin. By configuring appropriate Cache-Control headers or a CloudFront cache policy with optimal TTL values, you ensure that CloudFront caches the static images at edge locations for longer periods. This directly reduces the number of origin fetches, lowering data transfer costs without any changes to the application logic.

Exam trap

The trap here is that candidates may think disabling caching or bypassing CloudFront entirely will reduce costs, when in fact the opposite is true—effective caching is the key to reducing origin fetches and lowering data transfer costs.

Why the other options are wrong

B

Disabling CloudFront caching forces every request to the S3 origin, increasing origin fetches and data transfer costs, which is the opposite of the goal to reduce them.

D

A NAT Gateway is used to enable private subnets to access the internet or other AWS services, not to reduce bandwidth charges for CloudFront origins. It does not affect CloudFront caching or origin fetch behavior.

When would these options actually be correct?

B

A question asks for the most direct way to ensure users always see the latest version of an object without any delay, and cost is not a concern. Disabling caching would guarantee fresh content from S3 on every request.

D

In a scenario where an application in a private subnet needs to access an S3 bucket (or other internet resource) and you want to avoid using a public IP or an internet gateway, a NAT Gateway would be the correct solution to provide outbound internet access.

Why candidates pick the wrong answer

B

Candidates may think that disabling caching simplifies configuration or avoids stale content, but they overlook that caching is the primary mechanism to reduce origin load and costs.

D

Candidates may mistakenly think that a NAT Gateway can reduce data transfer costs because it is associated with network address translation and cost management, but it does not apply to CloudFront-to-S3 data transfer.

436
MCQhard

Based on the exhibit, the application tier is not replacing unhealthy instances even though the Auto Scaling group spans two Availability Zones. What change most directly improves automatic recovery when the application process fails?

A.Increase the ASG desired capacity so that extra instances absorb the failed ones.
B.Set the Auto Scaling group health check type to ELB so target group health determines replacement.
C.Replace the Application Load Balancer with a Network Load Balancer to improve failover speed.
D.Increase the HealthCheckGracePeriod to the maximum value so the instances have more time to stabilize.
AnswerB

This makes Auto Scaling replace instances that fail the load balancer health check even when EC2 status checks still pass. The exhibit shows the application health endpoint returns 500 while EC2 checks remain passing, so EC2-only health checks miss the failure. ELB-based health checks align replacement with real application availability.

Why this answer

Setting the Auto Scaling group health check type to ELB allows the ASG to use the target group's health checks, which monitor application-level health (e.g., HTTP 200 responses). When the application process fails, the ELB marks the instance as unhealthy, and the ASG immediately terminates and replaces it. This directly addresses the issue of unhealthy instances not being replaced, as the default EC2 health check only verifies instance status (e.g., running vs. stopped), not application responsiveness.

Exam trap

The trap here is that candidates assume the default EC2 health check is sufficient for application-level failures, but it only checks instance state (running/stopped), not the application process, so the ASG never triggers replacement for application crashes.

Why the other options are wrong

A

Increasing desired capacity adds more instances but does not fix the health check configuration; the ASG still uses EC2 status checks, which may not detect application-level failures, so unhealthy instances are not replaced.

C

The question is about replacing unhealthy instances based on application process failure, not about failover speed. An NLB does not provide application-level health checks, so it would not detect application process failures.

D

Increasing HealthCheckGracePeriod only delays the start of health checks, but does not fix the root cause: the ASG is using EC2 status checks (default) instead of ELB health checks, so it never detects application-level failures.

When would these options actually be correct?

A

In a scenario where the ASG is correctly configured to replace unhealthy instances but the application needs to handle sudden traffic spikes without downtime, increasing desired capacity ensures enough healthy instances are always available to absorb load.

C

When the requirement is to handle sudden traffic spikes with minimal latency and the application can tolerate connection-level health checks, or when the architecture needs to preserve the source IP address of clients.

D

This option would be correct in a scenario where instances are being prematurely terminated because health checks begin before the application finishes booting, causing false positives. Increasing the grace period gives the application more time to become healthy.

Why candidates pick the wrong answer

A

Candidates think that adding more instances provides redundancy, but they overlook that the core issue is the health check type not detecting application failures, so extra instances won't be triggered to replace unhealthy ones.

C

Candidates may think that an NLB's faster failover and lower latency would improve recovery, but they overlook that the issue is about detecting application-level health, which requires an ALB's HTTP health checks.

D

Candidates may think that giving instances more time to stabilize will prevent unnecessary replacements, but they overlook that the ASG must first be configured to use ELB health checks to detect application failures at all.

437
MCQmedium

A team runs an EC2-based service and ships logs to Amazon CloudWatch Logs. They enabled long log retention and turned on detailed monitoring to improve troubleshooting. Their monthly CloudWatch costs have grown unexpectedly. Compliance requires that the logs remain available in CloudWatch Logs (for querying and audits) for 90 days, and alerts/alarms do not require detailed EC2 monitoring. What change best reduces cost while meeting requirements?

A.Keep the current long retention and detailed monitoring; reduce the log volume by sampling 10% of events
B.Set the CloudWatch Logs retention to 90 days and disable detailed EC2 monitoring (use standard monitoring) for the instances
C.Move all logs to S3 immediately and delete the CloudWatch log groups to reduce costs
D.Increase CloudWatch alarm thresholds to reduce the number of metric datapoints
AnswerB

CloudWatch Logs storage costs are driven primarily by retention period. Setting retention to exactly 90 days reduces storage cost while meeting compliance. Disabling detailed EC2 monitoring reduces the number/granularity of metrics (detailed is billed more than standard), lowering monitoring cost without impacting alarms that don’t require high-resolution metrics.

Why this answer

Reduces costs by setting CloudWatch Logs retention to exactly 90 days (meeting compliance) and disabling detailed monitoring (which incurs per-minute metrics charges) in favor of standard 5-minute monitoring. This directly addresses the two main cost drivers—long retention and detailed EC2 monitoring—while preserving the required 90-day log availability for queries and audits.

Exam trap

The trap here is that candidates may think sampling logs or moving them to S3 is acceptable, but the requirement explicitly states logs must remain available in CloudWatch Logs for querying and audits, making those options non-compliant.

How to eliminate wrong answers

Option A is wrong because sampling only 10% of log events would lose critical data for troubleshooting and audits, violating the compliance requirement that logs remain available for 90 days. Option C is wrong because moving logs to S3 immediately and deleting CloudWatch log groups would remove the ability to query logs in CloudWatch Logs Insights, breaking the requirement that logs remain available in CloudWatch Logs for querying. Option D is wrong because increasing alarm thresholds does not reduce the number of metric datapoints collected; detailed monitoring still sends per-minute metrics, and thresholds only affect when alarms trigger, not the volume of data ingested or stored.

438
MCQmedium

A high-volume telemetry pipeline writes streaming click events that must be processed by multiple independent consumers. Which service is most appropriate? The architecture review board prefers a managed AWS-native control.

A.Amazon Kinesis Data Streams
B.AWS DataSync
C.Amazon EBS
D.Amazon Route 53
AnswerA

Amazon Kinesis Data Streams satisfies the multiple independent consumers constraint: each consumer reads the stream independently via its own iterator, without competing for messages, and retains data up to 365 days for replay. As a fully managed AWS-native service, it also meets the review board's managed control preference.

Why this answer

Amazon Kinesis Data Streams is the correct choice because it is a fully managed, AWS-native service designed for real-time streaming data ingestion and processing. It supports multiple independent consumers via enhanced fan-out, which provides each consumer with a dedicated throughput of up to 2 MB/sec per shard, ensuring that high-volume click events can be processed concurrently without contention.

Exam trap

The trap here is confusing batch data transfer services (DataSync) or storage services (EBS) with real-time streaming, leading candidates to overlook Kinesis Data Streams' native support for multiple independent consumers via enhanced fan-out.

How to eliminate wrong answers

Option B (AWS DataSync) is wrong because it is a data transfer service for moving large datasets between on-premises storage and AWS, not a real-time streaming pipeline. Option C (Amazon EBS) is wrong because it provides block-level storage volumes for EC2 instances, not a streaming data ingestion or processing capability. Option D (Amazon Route 53) is wrong because it is a DNS and domain name resolution service, completely unrelated to streaming telemetry data.

439
Multi-Selecthard

A payments API requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The architecture review board prefers a managed AWS-native control.

Select 2 answers
A.Deletion protection or tightly controlled delete permissions
B.Point-in-time recovery
C.Global secondary indexes
D.DAX
AnswersA, B

Deletion protection blocks table deletion outright, satisfying the accidental-delete requirement with a managed AWS-native control needing no custom tooling. Pair it with point-in-time recovery for continuous backups, enabling restore to any second within the 35-day window and meeting the low RPO.

Why this answer

Point-in-time recovery (PITR) enables continuous backups of the DynamoDB table, allowing restoration to any point within the last 35 days, which satisfies the requirement for point-in-time recovery. Deletion protection prevents accidental deletion of the table by blocking drop-table operations, meeting the accidental-delete protection requirement. Both are managed AWS-native controls that require no custom scripting or external tooling.

Exam trap

The trap here is that candidates often confuse operational features like DAX (caching) or GSIs (indexing) with data protection mechanisms, but neither provides backup/restore or deletion safeguards required for resilience and data durability.

440
MCQmedium

A trading dashboard uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered?

A.A single-AZ Aurora cluster
B.Aurora Global Database
C.Manual snapshots copied monthly
D.An ElastiCache Redis replica
AnswerB

Aurora Global Database replicates storage across Regions with typical lag under one second, giving low RPO and rapid promotion of a secondary Region to primary during disaster recovery. Single-Region Aurora replicas cannot deliver cross-Region failover at that recovery speed.

Why this answer

Aurora Global Database is designed for cross-Region disaster recovery with a typical RPO of 1 second and RTO of less than 1 minute, using storage-based replication that does not impact database performance. This meets the low RPO requirement for a trading dashboard, where data loss must be minimized.

Exam trap

The trap here is that candidates might choose manual snapshots (Option C) thinking they are sufficient for DR, but they overlook the critical requirement of low RPO, which snapshots copied monthly cannot satisfy.

How to eliminate wrong answers

Option A is wrong because a single-AZ Aurora cluster provides no cross-Region replication and offers no disaster recovery across AWS Regions, resulting in potentially high RPO if the primary Region fails. Option C is wrong because manual snapshots copied monthly have an RPO of up to one month, which is far too high for a trading dashboard requiring low RPO. Option D is wrong because ElastiCache Redis is an in-memory cache, not a persistent database, and cannot serve as a cross-Region disaster recovery solution for Aurora MySQL data.

441
MCQeasy

A startup stores application configuration files in an Amazon S3 bucket. The security team wants to ensure that objects in the bucket are encrypted at rest with keys that the company manages and can rotate on its own schedule. Which S3 encryption option should a solutions architect choose?

A.Server-side encryption with AWS KMS keys (SSE-KMS) using a customer managed key.
B.Server-side encryption with customer-provided keys (SSE-C).
C.Client-side encryption with an AWS KMS key before uploading objects to the bucket.
D.Server-side encryption with Amazon S3 managed keys (SSE-S3).
AnswerA

SSE-KMS with a customer managed key lets the company control the key policy, define who can use the key, and configure automatic key rotation on a schedule the company chooses. This meets the requirement for company-managed keys with controllable rotation, and it also provides an audit trail of key usage through AWS CloudTrail, which is valuable for compliance.

Why this answer

SSE-KMS with a customer managed key gives the company ownership of the key policy, the ability to enable and schedule automatic key rotation, and CloudTrail visibility into key usage. The other S3 encryption options either leave key management to AWS, require the application to supply keys per request, or move encryption entirely to the client, none of which matches the requirement for company-managed, rotatable keys.

Exam trap

The trap here is assuming SSE-S3 allows customer-controlled rotation, when SSE-S3 keys are fully managed by AWS and cannot be rotated on a customer-defined schedule.

442
MCQeasy

A startup runs a stateless web application on a single Amazon EC2 instance in one Availability Zone. The application has become popular, and the startup wants to ensure that the application can survive the failure of an Availability Zone and can handle increased traffic. Which architecture change should the startup make FIRST?

A.Resize the existing EC2 instance to a larger instance type to handle increased traffic and improve availability.
B.Create an Amazon Machine Image (AMI) of the instance and store it in Amazon S3 for quick recovery after a failure.
C.Enable detailed monitoring on the EC2 instance and create Amazon CloudWatch alarms to notify the team of failures.
D.Move the application to an Auto Scaling group that spans multiple Availability Zones, and place an Application Load Balancer in front of it.
AnswerD

An Auto Scaling group spanning multiple Availability Zones distributes instances across AZs, so the application survives an AZ failure. An Application Load Balancer distributes traffic and performs health checks, routing only to healthy instances. This provides both high availability and horizontal scalability, directly addressing the requirements with minimal changes.

Why this answer

Moving to an Auto Scaling group across multiple Availability Zones with an Application Load Balancer provides both high availability and elasticity. The load balancer health checks ensure traffic goes only to healthy instances, and the Auto Scaling group replaces failed instances automatically. This architecture eliminates single points of failure and allows the application to scale with demand, which is the fundamental first step for resilience.

Exam trap

The trap here is focusing on vertical scaling or monitoring as a way to achieve high availability, when the real solution requires distributing the workload across multiple Availability Zones.

443
MCQmedium

An events service publishes critical notifications using Amazon SNS. Three independent downstream systems (A, B, and C) subscribe to the topic. Downstream system B sometimes fails to process certain messages (for example, it times out or returns an error while handling the message), and you want: 1) failures in B to be isolated so A and C keep processing unaffected, and 2) messages that B cannot successfully process after retries to be sent to a DLQ for B. Which design best meets these requirements?

A.Subscribe each downstream directly with HTTPS endpoints and configure a single SNS dead-letter queue (DLQ) for the topic.
B.For each downstream system, create its own SQS queue, subscribe each SQS queue to the SNS topic, and configure a redrive policy with a DLQ for each SQS queue.
C.Use one shared SQS queue for all three downstream systems and configure a single DLQ only when all three downstream systems fail.
D.Use EventBridge rules to invoke A, B, and C synchronously with retries enabled, and send failures to a common DLQ.
AnswerB

SNS delivers the message independently to each subscribed SQS queue. If downstream B fails to process a message, B can avoid deleting it from its own queue; after visibility timeout and retry attempts, SQS redrives messages to B’s DLQ. A and C are isolated because they have separate queues and DLQs, so B’s failures do not prevent deliveries to A and C.

Why this answer

It creates a dedicated SQS queue for each downstream system, which isolates failures: if system B fails, its SQS queue will accumulate messages while systems A and C continue processing from their own queues. Each SQS queue can have a redrive policy that moves messages to a per-queue DLQ after the configured maximum retries are exhausted, satisfying the requirement for a B-specific DLQ without affecting the other subscribers.

Exam trap

The trap here is that candidates assume a single DLQ at the SNS topic level is sufficient, but SNS DLQs only apply to the SNS delivery failure (e.g., HTTP endpoint unreachable), not to downstream processing failures after the message is delivered to SQS.

How to eliminate wrong answers

Option A is wrong because a single SNS DLQ applies to the entire topic, not per-subscriber; if B fails, messages would be sent to the common DLQ for all subscribers, and A and C would still receive the message from SNS, but the DLQ is not isolated to B. Option C is wrong because a shared SQS queue for all three systems means a failure in B could block or delay messages for A and C, and a single DLQ would trigger only when all three fail, not when B alone fails. Option D is wrong because EventBridge synchronous invocation with a common DLQ would cause failures in B to potentially block or delay A and C (since synchronous calls are sequential), and the DLQ is shared, not isolated to B.

444
MCQmedium

A dev sandbox has unpredictable DynamoDB traffic with long idle periods and occasional spikes. Which capacity mode should minimize operational overhead and avoid paying for idle provisioned capacity? The architecture review board prefers a managed AWS-native control.

A.Reserved capacity for maximum daily traffic
B.Provisioned capacity set for peak traffic
C.DynamoDB on-demand capacity mode
D.Global tables in every Region
AnswerC

DynamoDB on-demand capacity mode charges only for the read and write requests you actually make, with no minimum capacity, no capacity planning, and no manual scaling. It can instantly scale from zero to any required throughput to absorb traffic spikes, making it ideal for unpredictable workloads like a development sandbox that may sit idle for hours or days. During idle periods your bill drops to near zero, and you never have to worry about throttle errors or forecasting demand—exactly matching the workload characteristics described in the question.

Why this answer

DynamoDB on-demand capacity mode (Option C) is ideal for unpredictable traffic with long idle periods and spikes because it automatically scales to handle workload demands without requiring any capacity planning. You pay only for the reads and writes you perform, eliminating the cost of idle provisioned capacity and the operational overhead of managing scaling thresholds.

Exam trap

The trap here is that candidates may confuse 'Reserved capacity' (an EC2/RDS concept) with DynamoDB pricing, or assume Provisioned capacity is always cheaper without considering the cost of idle resources in unpredictable workloads.

How to eliminate wrong answers

Option A is wrong because Reserved capacity is not a DynamoDB pricing model; it applies to Amazon RDS and EC2, not DynamoDB, and would lock you into a fixed cost regardless of usage. Option B is wrong because Provisioned capacity set for peak traffic would require you to pay for the peak capacity even during idle periods, leading to wasted cost and manual scaling adjustments. Option D is wrong because Global tables are a replication feature for multi-Region active-active setups, not a capacity mode; they add complexity and cost without addressing the need to avoid paying for idle provisioned capacity.

445
MCQmedium

Based on the exhibit, what is the most appropriate fix so the workload in Account A can access the S3 bucket in Account B without using long-lived access keys?

A.Create an IAM role in Account B, trust Account A's AppRole to assume it with STS, and then access the bucket using temporary credentials.
B.Attach AmazonS3FullAccess to the instance profile role in Account A and keep using the same direct access path.
C.Add an SCP to Account A that allows S3 actions against buckets in Account B.
D.Enable S3 versioning on the bucket so cross-account requests are automatically trusted.
AnswerA

Assuming a role in the target account is a clean cross-account pattern that uses temporary credentials instead of static keys. The trust policy in Account B controls who may assume the role, and the role in B can then be given the exact S3 permissions needed. This is easy to revoke centrally by changing the trust relationship or role policy.

Why this answer

It uses AWS Security Token Service (STS) to allow the workload in Account A to assume an IAM role in Account B, obtaining temporary credentials that grant access to the S3 bucket. This eliminates the need for long-lived access keys and follows the principle of least privilege, as the role can be scoped to specific S3 actions and resources.

Exam trap

The trap here is that candidates often confuse SCPs with resource-based policies or assume that attaching a managed policy to an instance profile automatically grants cross-account access, overlooking the need for explicit trust and bucket policies in the target account.

Why the other options are wrong

B

Option B is wrong because attaching AmazonS3FullAccess to the instance profile role in Account A does not grant cross-account access to an S3 bucket in Account B. The bucket's bucket policy must explicitly allow the role from Account A, and using long-lived access keys is not avoided.

C

SCPs (Service Control Policies) are used to restrict permissions in AWS Organizations accounts, not to grant cross-account access. They cannot allow actions; they only deny or limit permissions. Thus, adding an SCP to Account A does not enable the workload to access the S3 bucket in Account B.

D

Enabling S3 versioning does not grant cross-account access permissions; it only preserves object versions. Cross-account access requires explicit IAM policies and bucket policies, not versioning.

When would these options actually be correct?

B

Option B would be correct if the S3 bucket is in the same account as the workload (Account A) and the goal is to grant full S3 access to an EC2 instance using an instance profile role, without requiring cross-account access.

C

An SCP would be correct if the question asked: 'How can an organization administrator prevent all accounts in the organization from deleting an S3 bucket in a specific account?' In that case, adding an SCP that denies s3:DeleteBucket actions for that bucket would be appropriate.

D

A question asks how to protect against accidental deletion or overwriting of objects in an S3 bucket. Enabling versioning would be the correct answer because it allows recovery of previous object versions.

Why candidates pick the wrong answer

B

Candidates may think that attaching a full-access policy to the instance profile role is sufficient for any S3 access, overlooking the need for cross-account bucket policies and the requirement to avoid long-lived keys.

C

Candidates may confuse SCPs with resource-based policies or IAM policies, thinking SCPs can grant cross-account access, or they may overestimate the scope of SCPs as a general permission-granting mechanism.

D

Candidates may mistakenly think versioning enables cross-account trust or that it automatically resolves access control issues, confusing versioning with bucket policies or resource-based permissions.

446
MCQmedium

A company hosts a image sharing application on EC2. Administrators must connect without opening SSH or RDP ports to the internet. What should the architect use?

A.AWS Systems Manager Session Manager with the required instance role
B.An internet gateway attached to the private subnet
C.A public Elastic IP address on each instance
D.A bastion host with SSH open to 0.0.0.0/0
AnswerA

AWS Systems Manager Session Manager gives you encrypted, browser-based shell access to EC2 instances without opening any inbound ports. The instance must have the SSM agent installed and be assigned an IAM instance role with AmazonSSMManagedInstanceCore, allowing it to poll the SSM API over TLS. User access is governed by IAM policies and all shell activity is logged to CloudTrail and optionally S3/CloudWatch Logs, giving audited secure administration.

Why this answer

AWS Systems Manager Session Manager allows administrators to establish secure shell (SSH) or PowerShell (RDP) sessions to EC2 instances without opening any inbound ports. It uses the SSM Agent and the AWS Systems Manager service, which initiates outbound connections to the AWS cloud over HTTPS (port 443). The required instance role grants permissions for the agent to communicate with Systems Manager, enabling secure, auditable access without public IP addresses or bastion hosts.

Exam trap

The trap here is that candidates often default to a bastion host (Option D) as a traditional solution, but fail to recognize that a bastion host still requires opening SSH/RDP to the internet (even if only to the bastion), which violates the 'without opening SSH or RDP ports to the internet' constraint.

How to eliminate wrong answers

Option B is wrong because an internet gateway attached to a private subnet does not provide direct connectivity to instances; it only enables outbound internet access via a NAT device, and does not allow inbound administrative connections without opening ports. Option C is wrong because assigning a public Elastic IP address to each instance would expose them to the internet, requiring SSH or RDP ports to be open, which violates the requirement to not open those ports. Option D is wrong because a bastion host with SSH open to 0.0.0.0/0 exposes the bastion to the entire internet, creating a security risk and still requires opening SSH (port 22) to the internet, which directly contradicts the requirement.

447
MCQmedium

A stateless web API runs on EC2 instances behind an Application Load Balancer (ALB). The Auto Scaling group (ASG) currently uses subnets from only one Availability Zone, even though the ALB spans two Availability Zones. During maintenance of that single AZ, the ALB remains up but clients see timeouts because there are no healthy targets. Which change most directly improves resilience against an AZ failure?

A.Keep the ASG in one subnet/AZ, but enable ALB stickiness to reduce session interruption.
B.Update the ASG to launch instances across subnets in at least two Availability Zones and ensure ALB health checks target an application-ready path.
C.Add a NAT gateway in the public subnets so instances can reach the internet during maintenance events.
D.Create a second ALB in the same Availability Zone and route traffic using DNS failover.
AnswerB

Spreading the ASG across subnets in two Availability Zones removes the single-AZ failure domain, so the ALB can route to healthy targets in the surviving zone during maintenance. Pointing health checks at an application-ready path ensures instances register only when genuinely serving traffic, preventing the timeouts seen previously.

Why this answer

The most direct fix for AZ failure resilience is to distribute the ASG across multiple Availability Zones. With the ALB already spanning two AZs, if the ASG only launches instances in one AZ, a failure of that AZ leaves the ALB with zero healthy targets, causing timeouts. By configuring the ASG to launch instances in at least two AZs and setting ALB health checks to an application-ready path, the ALB can route traffic to healthy instances in the surviving AZ, maintaining availability.

Exam trap

The trap here is that candidates may think adding a second ALB or enabling stickiness solves the problem, when the real issue is that the ASG is not distributing instances across multiple Availability Zones, leaving the ALB with no healthy targets during an AZ outage.

Why the other options are wrong

A

Enabling ALB stickiness does not address the root cause: the ASG has no healthy targets in the only AZ, so the ALB cannot route traffic to any instance, causing timeouts regardless of stickiness.

C

The issue is that EC2 instances are only in one AZ, so adding a NAT gateway does not provide healthy targets in the other AZ; NAT gateways enable outbound internet access but do not affect ALB target availability.

D

Creating a second ALB in the same AZ does not address the lack of healthy targets in the other AZ; the ALB is already spanning two AZs, but the ASG only has instances in one AZ. DNS failover would still route to the same AZ if both ALBs are in the same AZ, failing to provide resilience against an AZ failure.

When would these options actually be correct?

A

This option would be correct in a scenario where the application is stateful and sessions are stored locally on instances, and the goal is to maintain session persistence to avoid data loss during normal operations (not AZ failure).

C

This option would be correct if the question described private subnets without internet access, and the EC2 instances needed to download updates or access external APIs for the application to function properly.

D

This option would be correct in a scenario where the ALB is only in one AZ and you need to achieve cross-AZ failover for the load balancer itself. For example, if the ALB is deployed in a single AZ and you want to ensure high availability by placing a second ALB in another AZ with DNS failover using Route 53.

Why candidates pick the wrong answer

A

Candidates may think stickiness helps during failures by keeping users on the same instance, but it does not solve the problem of having zero healthy targets in the only AZ.

C

Candidates may confuse network connectivity (NAT) with high availability, thinking that internet access is required for the ALB to route traffic to instances, or they may misattribute the timeout issue to a lack of outbound connectivity.

D

Candidates may think that adding a second ALB provides redundancy, but they overlook that the root cause is the ASG's single-AZ deployment, not the ALB's availability. The ALB already spans two AZs, so the issue is the lack of targets in the second AZ.

448
MCQmedium

An orders service publishes payment instructions to an Amazon SQS Standard queue. A downstream consumer sometimes times out or crashes after it has partially completed processing, causing the same instruction to be processed more than once. You must keep the design resilient without attempting to guarantee exactly-once processing. Which approach best handles duplicates safely?

A.Set the SQS visibility timeout extremely long so the message cannot be retried even after processing failures.
B.Make the consumer idempotent by deriving a deterministic idempotency key from the payment instruction (for example, the instruction ID), persisting the result of successful processing, and skipping re-processing when that key is already marked successful.
C.Switch to an SQS FIFO queue but remove error handling in the consumer so duplicates never occur.
D.Send all failed messages to a DLQ and rely on it to deduplicate messages that were already successfully processed.
AnswerB

SQS Standard provides at-least-once delivery, so duplicates are expected. Idempotency ensures that re-processing the same instruction does not create incorrect side effects. Persisting a deterministic key/result allows the consumer to safely short-circuit duplicates after retries/timeouts.

Why this answer

Making the consumer idempotent ensures that even if the same payment instruction is processed multiple times due to timeouts or crashes, the system remains consistent. By deriving a deterministic idempotency key (e.g., the instruction ID) and persisting the result of successful processing, the consumer can skip re-processing when the key is already marked as successful. This approach aligns with the requirement to keep the design resilient without guaranteeing exactly-once processing, as it safely handles duplicates at the application level.

Exam trap

The trap here is that candidates often assume SQS FIFO queues or DLQs inherently solve duplicate processing, but the exam tests understanding that Standard queues require application-level idempotency for safe duplicate handling, and FIFO queues do not eliminate the need for idempotent consumers in crash scenarios.

How to eliminate wrong answers

Option A is wrong because setting the SQS visibility timeout extremely long does not prevent duplicates; it only delays retries, and if the consumer crashes after partially processing, the message will eventually become visible again and be reprocessed, leading to duplicates. Option C is wrong because switching to an SQS FIFO queue provides exactly-once processing within a five-minute deduplication window, but removing error handling in the consumer does not prevent duplicates from timeouts or crashes; FIFO queues still allow retries, and without error handling, the system becomes fragile. Option D is wrong because a Dead-Letter Queue (DLQ) is used to capture messages that fail after multiple retries, not to deduplicate messages; relying on a DLQ for deduplication is a misconception, as DLQs do not track successful processing and cannot prevent duplicates from being processed again.

449
Multi-Selecthard

A reporting application in Account B must read files from an S3 bucket in Account A. The bucket contains objects encrypted with a customer managed KMS key in Account A. The application role in Account B already has an identity policy allowing s3:GetObject on the bucket prefix, but requests still fail with AccessDenied. Which two changes are required for the application to read the objects? Select two.

Select 2 answers
A.Add a bucket policy in Account A that allows the Account B role to perform s3:GetObject on the required prefix.
B.Add the Account B role to the KMS key policy in Account A with permission to use kms:Decrypt.
C.Attach an IAM policy in Account B that grants s3:* on the bucket and its objects.
D.Create an S3 gateway endpoint in Account B so the application can reach the bucket privately.
E.Add an SCP in Account A that allows the Account B role to bypass KMS encryption checks.
AnswersA, B

Cross-account S3 access requires a resource-based permission on the bucket. The bucket policy must explicitly allow the external role to read the needed prefix, otherwise the bucket owner blocks the request even if the role's identity policy allows it.

Why this answer

Cross-account S3 access requires the destination account (Account A) to explicitly grant access via a bucket policy that allows the source account's role (Account B) to perform s3:GetObject on the specified prefix. Without this bucket policy, the S3 service in Account A will deny the request, even if the IAM identity policy in Account B permits the action. Option B is correct because the objects are encrypted with a customer managed KMS key in Account A; the application role in Account B must be added to the KMS key policy with kms:Decrypt permission to decrypt the objects during retrieval.

Both the S3 bucket policy and the KMS key policy are required for cross-account encrypted access.

Exam trap

The trap here is that candidates often assume a cross-account IAM role with s3:GetObject permission is sufficient, overlooking that S3 bucket policies and KMS key policies are separate authorization layers that must explicitly allow the external principal, especially when objects are encrypted with a customer managed KMS key.

Why the other options are wrong

C

The application role in Account B already has an identity policy allowing s3:GetObject on the bucket prefix, so adding another IAM policy granting s3:* is redundant and does not address the cross-account permission issue or the KMS key policy requirement.

D

The error is AccessDenied, not connectivity issues. S3 Gateway Endpoints only provide private network access to S3, but do not grant IAM permissions or resolve KMS decryption authorization failures.

E

SCPs (Service Control Policies) cannot grant permissions; they only restrict permissions. Additionally, SCPs cannot bypass KMS encryption checks; the KMS key policy must explicitly allow the Account B role to use kms:Decrypt.

When would these options actually be correct?

C

This option would be correct if the application role in Account B did not already have the necessary S3 permissions, and the question asked for a single-account setup where the bucket and role are in the same account, requiring only IAM policy updates.

D

A question where an application in a VPC cannot reach S3 due to public internet routing restrictions (e.g., no NAT gateway, no internet gateway) and the bucket is in the same region. Adding an S3 Gateway Endpoint would provide private connectivity.

E

An SCP would be correct in a question where an organization wants to prevent all accounts from disabling encryption on S3 buckets, and the correct action is to add an SCP that denies s3:PutBucketEncryption or similar actions to enforce encryption requirements.

Why candidates pick the wrong answer

C

Candidates may think that granting broader S3 permissions (s3:*) will override any missing permissions, not realizing the core issue is cross-account access and KMS key policy, not insufficient IAM permissions in Account B.

D

Candidates may confuse network connectivity problems with authorization failures, assuming that a private endpoint can bypass IAM or KMS permission issues.

E

Candidates may confuse SCPs with resource policies or think SCPs can grant cross-account access, or they may misunderstand that SCPs can override KMS permissions, leading them to select this option as a shortcut to fix the access issue.

450
MCQmedium

A marketing site runs on x86 EC2 instances and uses open-source software with no architecture-specific licensing restriction. What should be evaluated to reduce compute cost? The design must avoid adding custom operational scripts.

A.Cross-Region data replication for all data
B.io2 Block Express volumes for all instances
C.AWS Graviton-based instances after performance testing
D.Dedicated Hosts by default
AnswerC

Graviton instances often provide better price performance for compatible workloads.

Why this answer

AWS Graviton-based instances (ARM architecture) offer up to 40% better price-performance compared to comparable x86 instances for many workloads. Since the marketing site uses open-source software with no architecture-specific licensing restrictions, migrating to Graviton after performance testing can significantly reduce compute costs without requiring custom operational scripts, as the OS and software can be recompiled for ARM natively.

Exam trap

The trap here is that candidates may assume Dedicated Hosts (Option D) are a cost-saving measure, but they actually increase costs unless you have specific licensing needs, and they violate the 'no custom operational scripts' constraint by requiring manual host management.

How to eliminate wrong answers

Option A is wrong because Cross-Region data replication increases data transfer and storage costs, and it does not reduce compute costs; it is a disaster recovery or latency optimization strategy, not a cost-saving measure for compute. Option B is wrong because io2 Block Express volumes are high-performance, high-cost SSD volumes designed for latency-sensitive workloads like databases, not for reducing compute costs; they would increase storage costs without affecting compute efficiency. Option D is wrong because Dedicated Hosts are a licensing option that incurs additional per-host charges and are only cost-effective for specific scenarios like bring-your-own-license (BYOL) software with socket/core restrictions; they do not reduce compute costs for open-source software and would increase operational overhead.

Page 5

Page 6 of 13

Page 7