Courseiva

SAA-C03 (SAA-C03) — Questions 226–300

935 questions total · 13pages · All types, answers revealed

Page 3

Page 4 of 13

Page 5
226
MCQmedium

A company stores RDS database credentials in AWS Systems Manager Parameter Store as SecureString parameters. The security team requires that database passwords rotate automatically every 30 days. Which change should a solutions architect recommend?

A.Create a scheduled EventBridge rule to invoke a Lambda function that updates the Parameter Store SecureString value every 30 days
B.Migrate the credentials to AWS Secrets Manager and enable automatic rotation with a 30-day schedule
C.Enable Parameter Store SecureString automatic rotation in the AWS console
D.Configure AWS Config to detect password age and trigger an SNS notification after 30 days
AnswerB

AWS Secrets Manager is the only service in the options that natively integrates with Amazon RDS to rotate credentials with a managed Lambda rotation function, which you can configure to run every 30 days. The rotation process updates the database password, the secret value, and the secret's versions atomically and can be tested for rollback, eliminating the need for custom rotation code. Enabling rotation adds minimal operational overhead — you just choose the 30-day interval and the rotation Lambda provisions itself with the appropriate IAM role and permissions. This directly meets the requirement and is the recommended AWS pattern for automated RDS credential rotation.

Why this answer

AWS Secrets Manager provides native automatic rotation for RDS credentials using a managed Lambda function that rotates the secret on a defined schedule and updates the database password atomically.

Parameter Store SecureString does not support built-in automatic rotation — rotation must be implemented manually with custom automation. Secrets Manager is specifically designed for secrets requiring lifecycle management including rotation, auditing, and fine-grained access control.

Exam trap

Both services encrypt values using KMS, which causes candidates to treat them as equivalent. Only Secrets Manager provides automatic rotation with managed Lambda integration and rotation history. Parameter Store is appropriate for configuration values and static secrets.

Whenever automatic rotation is a security policy requirement, Secrets Manager is the answer.

Why the other options are wrong

A

Creating a custom EventBridge rule + Lambda for rotation works but requires development and maintenance effort. It lacks native rotation history and is more complex than the purpose-built Secrets Manager solution.

C

There is no built-in automatic rotation toggle in Parameter Store. This feature does not exist in the Parameter Store console — automatic rotation is a Secrets Manager capability.

D

AWS Config detects and alerts on compliance drift but cannot automatically rotate a database password. SNS notification is a detection mechanism, not a remediation mechanism.

227
MCQhard

Based on the exhibit, an application in the same AWS account can upload and read objects in an S3 bucket encrypted with a customer managed KMS key, but GetObject fails with an AccessDenied error from AWS KMS. The IAM role already has s3:GetObject, s3:PutObject, kms:Decrypt, and kms:GenerateDataKey permissions. What change most directly fixes the issue while preserving least privilege?

A.Add an S3 bucket ACL that grants the application role full control over objects.
B.Update the KMS key policy to allow the application role to use the key, ideally with a kms:ViaService condition for S3.
C.Replace the customer managed key with the AWS managed S3 key so IAM permissions become sufficient.
D.Add an S3 bucket policy that grants s3:GetObject and s3:PutObject to the role for all objects.
AnswerB

To resolve a KMS key policy denial for an application using S3 with a customer managed key, the KMS key policy must explicitly list the application's IAM role as a principal allowed to call kms:Decrypt (and kms:GenerateDataKey for writes). IAM permissions alone are not enough because KMS key policies act as a separate authorization layer; the key policy must explicitly trust the principal. Adding a kms:ViaService condition set to "s3.amazonaws.com" enforces least privilege by restricting key usage to requests that originate through S3, preventing the role from using the key outside the intended service context.

Why this answer

The error is an AccessDenied from AWS KMS, not from S3, which means the IAM role has the required S3 permissions (s3:GetObject) and KMS API permissions (kms:Decrypt), but the KMS key policy does not explicitly grant the role access to the key. Since customer managed KMS keys require a key policy to grant IAM principals permission to use the key (IAM policies alone are insufficient unless the key policy delegates such authority), updating the key policy to allow the role with a kms:ViaService condition for S3 directly resolves the KMS-side denial while preserving least privilege.

Exam trap

The trap here is that candidates see 'AccessDenied' and assume the S3 bucket policy or ACL is missing, when the error message explicitly states it is from AWS KMS, meaning the fix must be at the KMS key policy level, not the S3 resource policy.

How to eliminate wrong answers

Option A is wrong because S3 bucket ACLs control access to S3 objects themselves, not KMS key permissions; the error is from KMS, not S3, so an ACL cannot fix a KMS AccessDenied. Option C is wrong because switching to the AWS managed S3 key (SSE-S3) would remove the need for KMS permissions entirely, but it changes the encryption type and does not preserve the use of a customer managed key as required by the scenario; it also violates least privilege by removing control over the key. Option D is wrong because the IAM role already has s3:GetObject and s3:PutObject permissions, and the error is from KMS, not S3; adding a bucket policy for the same S3 actions does not address the missing KMS key policy grant.

228
MCQmedium

A financial services company runs a three-tier web application on AWS. The application tier consists of Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. A security audit reveals that the application instances are receiving large volumes of unwanted traffic directly from the internet on port 443, bypassing the load balancer. The company wants to ensure that only traffic from the ALB can reach the application instances, while allowing the instances to download software updates from the internet. What should a solutions architect recommend?

A.Modify the network ACL on the application subnets to deny inbound traffic on port 443 from all sources except the ALB's private IP addresses.
B.Attach an AWS WAF web ACL to the ALB and create a rule to block all IP addresses except those in the ALB's subnet CIDR range.
C.Move the application instances to a placement group and enable enhanced networking to prevent direct internet access.
D.Configure the application instances' security group to allow inbound traffic only from the ALB's security group, and place the instances in private subnets with a NAT gateway for outbound internet access.
AnswerD

Referencing the ALB's security group as the source in the instances' inbound rule ensures only traffic that passed through the load balancer is accepted. Placing instances in private subnets removes direct internet routing, and a NAT gateway provides outbound-only internet access for updates, satisfying both requirements without exposing the instances.

Why this answer

The most secure and operationally sound approach is to use security group referencing so that the instances accept traffic only from the load balancer, and to remove direct internet exposure by placing instances in private subnets. A NAT gateway then allows outbound updates. This combination enforces the traffic path through the ALB while preserving necessary outbound connectivity.

Exam trap

The trap here is assuming that an AWS WAF rule or a network ACL can restrict traffic that reaches instances directly, when only security group referencing combined with private subnets removes the direct path.

229
MCQmedium

A log archive serves infrequently accessed user documents that must be available immediately when requested. Which S3 storage class is likely the best cost fit?

A.Instance store volumes
B.S3 Standard-IA or S3 One Zone-IA depending on resilience requirements
C.S3 Standard for all objects
D.S3 Glacier Deep Archive
AnswerB

S3 Standard-IA and One Zone-IA both provide immediate millisecond retrieval, satisfying the instant-access constraint, while their lower storage price suits infrequent access. Standard-IA replicates across three Availability Zones; One Zone-IA stores data in a single AZ, so choose it only when reduced resilience is acceptable.

Why this answer

S3 Standard-IA or S3 One Zone-IA is the best cost fit because the workload involves infrequently accessed data that requires immediate retrieval (millisecond latency). Standard-IA offers lower storage cost than S3 Standard while maintaining high durability and low-latency access, and One Zone-IA provides even lower cost for data that can tolerate a single-AZ failure. Both classes meet the 'available immediately' requirement, unlike Glacier tiers which have retrieval delays.

Exam trap

The trap here is that candidates often confuse 'infrequently accessed' with 'archival' and choose Glacier Deep Archive, forgetting that the requirement for immediate availability eliminates any Glacier tier due to its retrieval delays.

How to eliminate wrong answers

Option A is wrong because instance store volumes are ephemeral block storage attached to EC2 instances, not an S3 storage class, and they lose data on instance stop/termination, making them unsuitable for durable log archives. Option C is wrong because S3 Standard is designed for frequently accessed data with higher storage cost per GB, leading to unnecessary expense for infrequently accessed logs. Option D is wrong because S3 Glacier Deep Archive has retrieval times of 12–48 hours, which violates the 'available immediately' requirement.

230
MCQeasy

A company stores user uploads in an S3 bucket. Objects are accessed rarely after upload, but when an object is accessed, it must be retrievable quickly (minutes to a few hours). Objects must be retained for at least 18 months. The team wants to reduce storage cost while meeting these requirements. Which lifecycle configuration best fits these requirements?

A.Keep all objects in S3 Standard permanently to avoid lifecycle transition fees.
B.After 30 days, transition objects to S3 Glacier Instant Retrieval, and after 18 months, expire (delete) the objects.
C.After 30 days, transition objects to S3 Intelligent-Tiering, and set expiration to 12 months.
D.After 30 days, transition objects to S3 Glacier Deep Archive, and set expiration to 18 months.
AnswerB

The prompt requires (1) cost reduction for data that becomes infrequently accessed and (2) quick retrieval when accessed again, and (3) a minimum retention of at least 18 months. Glacier Instant Retrieval is intended for data that is accessed occasionally and needs fast retrieval. Transitioning after 30 days moves the long-term, rarely accessed portion of the data to a cheaper class, while expiring at 18 months satisfies the explicit retention requirement (the objects remain for at least 18 months).

Why this answer

It transitions objects to S3 Glacier Instant Retrieval after 30 days, which provides millisecond retrieval for rarely accessed data, meeting the quick retrieval requirement. The 18-month expiration ensures compliance with the retention policy while minimizing storage costs compared to keeping data in S3 Standard.

Exam trap

The trap here is that candidates may confuse retrieval time requirements: S3 Glacier Deep Archive is cheaper but has retrieval times of hours, not minutes, and S3 Intelligent-Tiering is for unpredictable access, not for data that is rarely accessed after upload.

How to eliminate wrong answers

Option A is wrong because keeping all objects in S3 Standard permanently ignores the cost-saving opportunity of lifecycle transitions; S3 Standard is more expensive for rarely accessed data, and there are no lifecycle transition fees for moving to colder storage classes. Option C is wrong because S3 Intelligent-Tiering is designed for unpredictable access patterns, not for data that is rarely accessed after upload, and setting expiration to 12 months violates the 18-month retention requirement. Option D is wrong because S3 Glacier Deep Archive has retrieval times of 12-48 hours, which does not meet the requirement of retrievable within minutes to a few hours.

231
MCQeasy

An organization hosts the same public API in two AWS Regions. Normal traffic should go to the primary Region. If the primary endpoint becomes unhealthy, Route 53 should automatically route users to the secondary Region. What is the best Route 53 configuration approach?

A.Use simple routing with one record that contains both regions as weighted targets.
B.Use weighted routing and set the secondary Region weight to 0 until needed.
C.Use Route 53 failover routing with health checks that mark the primary as unhealthy and fail over to the secondary.
D.Use latency-based routing so requests go to the region with the lowest latency, regardless of health.
AnswerC

Failover routing is designed for active/passive disaster recovery. You configure a primary record and a secondary record, each associated with health checks. When the primary fails its health checks, Route 53 automatically resolves the name to the secondary target.

Why this answer

Route 53 failover routing is designed for active-passive configurations where traffic is directed to a primary resource unless a health check marks it as unhealthy, at which point all traffic automatically shifts to the secondary resource. This directly matches the requirement of routing normal traffic to the primary Region and failing over to the secondary Region only when the primary endpoint becomes unhealthy.

Exam trap

The trap here is that candidates often confuse weighted routing with failover routing, mistakenly thinking that setting a weight of 0 on the secondary is a valid way to keep it inactive until needed, but Route 53 does not automatically adjust weights based on health checks.

How to eliminate wrong answers

Option A is wrong because simple routing does not support health checks or automatic failover; it simply returns all IP addresses in a random order, which cannot enforce a primary-secondary failover pattern. Option B is wrong because setting the secondary Region weight to 0 would prevent any traffic from reaching it even during a failure, and manually changing weights defeats the purpose of automatic failover. Option D is wrong because latency-based routing selects the Region with the lowest latency for each user, which does not guarantee that the primary Region handles normal traffic and does not automatically fail over based on endpoint health.

232
MCQmedium

Developers for a e-learning platform need temporary elevated access to production resources for troubleshooting. The security team wants approvals, expiry, and audit logging. Which approach is best?

A.Disable CloudTrail during troubleshooting
B.Use IAM Identity Center permission sets with time-bound access processes and CloudTrail auditing
C.Attach AdministratorAccess permanently to every developer role
D.Create shared administrator access keys for the team
AnswerB

Federated access with permission sets and audited temporary assignments reduces standing privilege.

Why this answer

IAM Identity Center permission sets allow you to define fine-grained permissions and assign them to users or groups with time-bound access (e.g., using a session duration or approval workflow). Combined with CloudTrail, every API call made during the elevated session is logged for audit, meeting the security team's requirements for approvals, expiry, and audit logging.

Exam trap

The trap here is that candidates may think IAM roles with a trust policy and temporary credentials are sufficient, but they overlook that IAM Identity Center provides centralized, time-bound permission sets with built-in approval workflows and audit integration, which is the best fit for the given requirements.

How to eliminate wrong answers

Option A is wrong because disabling CloudTrail during troubleshooting would eliminate audit logging, directly violating the security team's requirement for audit logging. Option C is wrong because permanently attaching AdministratorAccess to every developer role grants unrestricted, persistent elevated access with no expiry or approval process, violating the principle of least privilege and the need for time-bound access. Option D is wrong because creating shared administrator access keys for the team removes individual accountability, prevents proper audit trails (as actions cannot be attributed to a specific user), and provides no expiry or approval mechanism.

233
MCQmedium

A dev sandbox runs for several hours each night and can be interrupted and restarted. Which EC2 purchasing option should minimize cost?

A.On-Demand Instances only
B.Spot Instances
C.Dedicated Hosts
D.Provisioned IOPS volumes
AnswerB

Spot Instances offer spare AWS compute capacity at discounts of up to 90% compared to On-Demand, and they can be reclaimed by AWS with a two-minute warning when capacity is needed elsewhere. Because the dev sandbox runs only at night and can be interrupted without negative impact, it is an ideal fit for Spot. The main risk is interruption, but that is acceptable for this use case.

Why this answer

Spot Instances can be interrupted and restarted, making them ideal for fault-tolerant workloads like a nightly dev sandbox. They offer significant cost savings (up to 90% off On-Demand) because they use spare AWS EC2 capacity, which aligns perfectly with the scenario's tolerance for interruption.

Exam trap

The trap here is that candidates confuse 'interruptible' with 'unreliable' and choose On-Demand for stability, missing that Spot Instances are explicitly designed for fault-tolerant, non-critical workloads like a nightly dev sandbox.

How to eliminate wrong answers

Option A is wrong because On-Demand Instances are billed per second with no interruption, which is unnecessary for a workload that can be stopped and resumed, leading to higher costs. Option C is wrong because Dedicated Hosts provide physical servers for licensing or compliance needs, which is overkill and expensive for a simple dev sandbox. Option D is wrong because Provisioned IOPS volumes are a storage option (EBS), not an EC2 purchasing option, and do not directly affect compute cost optimization.

234
Multi-Selecthard

A third-party payroll vendor in another AWS account must assume a role in your account to write a daily settlement file to Amazon S3. You want to prevent confused-deputy attacks and make every assumed session traceable in CloudTrail back to an individual vendor user. Which three trust-policy or session controls should be used? Select three.

Select 3 answers
A.Specify the exact vendor role ARN as the trusted principal in the role trust policy.
B.Require an external ID in the trust policy conditions.
C.Require sts:SourceIdentity when the vendor assumes the role.
D.Use a wildcard principal and rely on the S3 bucket policy to narrow access later.
E.Give the vendor long-term IAM user credentials in your account for easier auditing.
AnswersA, B, C

The trust policy should name only the specific vendor role that is allowed to assume the role in your account. Restricting the principal minimizes the trust boundary and prevents unrelated identities from attempting the assumption path.

Why this answer

Specifying the exact vendor role ARN as the trusted principal in the trust policy ensures that only that specific role in the vendor's account can assume the role, preventing any other entity from impersonating the vendor. This is a key control to limit the trust boundary and avoid confused-deputy attacks.

Exam trap

The trap here is that candidates often think a bucket policy alone can control role assumption, but it cannot—the trust policy is the only mechanism to restrict which external principals can assume a role, and confused-deputy protections require explicit conditions like external ID and source identity.

Why the other options are wrong

D

Using a wildcard principal in the trust policy would allow any AWS principal to assume the role, violating the principle of least privilege and failing to prevent confused-deputy attacks. The S3 bucket policy cannot restrict who assumes the role, only what the assumed role can access.

E

Option E suggests giving the vendor long-term IAM user credentials in your account, which violates the principle of least privilege and makes auditing harder because actions are tied to a shared credential rather than individual vendor users. It also does not prevent confused-deputy attacks or ensure traceability to individual vendor users.

When would these options actually be correct?

D

In a scenario where you want to allow multiple accounts or services to assume a role without specifying each ARN individually, and you have additional controls like an external ID and source identity to prevent confused-deputy attacks, a wildcard principal might be acceptable if combined with strong condition keys.

E

A question where a trusted third party needs direct access to your AWS resources without assuming a role, and you have full control over their access policies. For example, a contractor who needs to upload files to S3 and you want to manage their permissions directly within your account, with CloudTrail logging tied to that IAM user.

Why candidates pick the wrong answer

D

Candidates may think that a bucket policy can compensate for a permissive trust policy, not realizing that the trust policy controls who can assume the role, while the bucket policy only controls actions after the role is assumed.

E

Candidates may think that giving the vendor their own IAM user in the account simplifies auditing because the user is directly visible in CloudTrail, but they overlook the security risks of sharing long-term credentials and the inability to trace actions back to individual vendor employees.

235
MCQmedium

Based on the exhibit, the web application must remain available even if one Availability Zone fails. What is the best change to improve resilience with the least redesign?

A.Increase DesiredCapacity to 4 while keeping all instances in subnet-a1.
B.Add subnet-b1 in a different Availability Zone to the Auto Scaling group.
C.Replace the Application Load Balancer with a Network Load Balancer.
D.Enable EBS encryption on the launch template volumes.
AnswerB

This spreads EC2 instances across two Availability Zones, so the Auto Scaling group can continue serving traffic if one AZ becomes unavailable. Because the ALB is already deployed in both subnets, this is the smallest change that adds true zonal resilience to the compute tier.

Why this answer

Adding subnet-b1 in a different Availability Zone to the Auto Scaling group ensures that EC2 instances are launched across two Availability Zones. If one zone fails, the ALB can route traffic to healthy instances in the other zone, maintaining application availability. This change requires minimal redesign because it only modifies the Auto Scaling group's subnet configuration without altering the load balancer or compute architecture.

Exam trap

The trap here is that candidates may think increasing instance count or changing load balancer type improves resilience, but without multi-AZ distribution, a single AZ failure still causes a total outage.

Why the other options are wrong

A

Increasing DesiredCapacity to 4 in a single subnet (subnet-a1) does not add resilience across Availability Zones; all instances remain in one AZ, so a failure of that AZ still causes total outage.

C

Replacing the Application Load Balancer with a Network Load Balancer does not improve resilience across Availability Zones; it only changes the load balancer type, which operates at a different layer and does not address the single-AZ failure risk.

D

Enabling EBS encryption does not improve availability or resilience across Availability Zones; it only protects data at rest. The question requires resilience against an AZ failure, which encryption does not address.

When would these options actually be correct?

A

This option would be correct if the question asked for a way to handle increased traffic while keeping all instances in the same subnet, and resilience across AZs was not required.

C

This option would be correct if the question required handling millions of requests per second with low latency, or if the application needed to preserve the source IP address for backend processing, as NLB operates at Layer 4 and supports these use cases.

D

This would be correct in a scenario where the question asks for a security improvement to protect sensitive data on EBS volumes, with no requirement for high availability or multi-AZ resilience.

Why candidates pick the wrong answer

A

Candidates may think that simply adding more instances provides high availability, overlooking that they must be distributed across multiple Availability Zones to survive an AZ failure.

C

Candidates may think that a Network Load Balancer is inherently more resilient or that changing load balancer types can fix availability issues, but resilience depends on multi-AZ deployment, not the load balancer type.

D

Candidates may confuse security measures with availability measures, or think that encryption adds redundancy, but encryption has no impact on fault tolerance.

236
MCQmedium

An API team runs an AWS Lambda function behind an Application Load Balancer (ALB). During predictable hourly traffic spikes, p95 response latency increases due to occasional cold starts. The team wants stable latency during those spikes without permanently overprovisioning resources for all functions. Which configuration is the most appropriate way to reduce cold starts for this Lambda function?

A.Publish a version of the function and configure provisioned concurrency on an alias, using autoscaling for the alias.
B.Increase the function memory size and rely on faster initialization to reduce cold starts.
C.Set reserved concurrency equal to the expected peak requests per second for the function.
D.Use an event source mapping with a higher batch size so Lambda triggers earlier and keeps the runtime warm.
AnswerA

Provisioned concurrency pre-initializes execution environments for a specific published function version. By attaching provisioned concurrency to an alias, you can control warm capacity and (with the right settings) autoscale the provisioned capacity for predictable spike patterns, reducing cold-start-driven latency increases.

Why this answer

Provisioned concurrency initializes a specified number of execution environments in advance, keeping them warm and ready to handle requests without cold start latency. By configuring provisioned concurrency on an alias with autoscaling, the team can dynamically adjust the number of pre-warmed environments to match predictable traffic spikes, avoiding permanent overprovisioning while ensuring stable p95 latency.

Exam trap

The trap here is confusing reserved concurrency (which limits concurrency but does not prevent cold starts) with provisioned concurrency (which pre-warms environments), leading candidates to select Option C as a cost-saving measure that fails to address latency.

Why the other options are wrong

B

Increasing memory size improves CPU speed and can reduce initialization time, but it does not eliminate cold starts; it only shortens them. The question asks to reduce cold starts, not just mitigate their duration.

C

Reserved concurrency limits the maximum number of concurrent executions but does not pre-warm instances; cold starts still occur during traffic spikes when new execution environments are needed.

D

Event source mappings (e.g., SQS, DynamoDB Streams) are not used with ALB triggers; ALB invokes Lambda synchronously via a function URL or alias ARN. Increasing batch size does not apply to ALB-triggered functions and does not prevent cold starts during predictable spikes.

When would these options actually be correct?

B

A question where the goal is to reduce the duration of cold starts (e.g., 'Which configuration minimizes the impact of cold starts on latency?') and the candidate is constrained from using provisioned concurrency due to cost or other restrictions.

C

In a scenario where a Lambda function must not exceed a specific concurrency limit to avoid throttling downstream resources (e.g., a database with connection limits), and the goal is to cap concurrent invocations rather than reduce latency.

D

For a Lambda function processing messages from an SQS queue or DynamoDB Streams, increasing the batch size can reduce the number of invocations and keep the runtime warm by processing more records per invocation, thereby mitigating cold starts during traffic spikes.

Why candidates pick the wrong answer

B

Candidates know that more memory often means faster Lambda execution, and they may incorrectly assume that faster initialization eliminates cold starts entirely, rather than just shortening them.

C

Candidates may confuse reserved concurrency with provisioned concurrency, thinking that reserving capacity eliminates cold starts, but reserved concurrency only guarantees available capacity, not pre-initialized environments.

D

Candidates may confuse event source mappings with ALB triggers and think that batching can pre-warm functions, not realizing that ALB invokes Lambda synchronously without batch processing.

237
MCQhard

A logistics firm runs an order-processing service that reads from an Amazon SQS queue and writes results to an Amazon DynamoDB table. During a marketing event, the consumer fleet scaled out aggressively and DynamoDB began returning ProvisionedThroughputExceededException errors, causing messages to be retried and some orders to be processed twice. The architects want to absorb traffic spikes without overprovisioning capacity and without duplicate processing. Which combination of changes should they make?

A.Add a global secondary index on the message ID and query it before each write to check whether the order has already been stored.
B.Switch the DynamoDB table to on-demand capacity mode and have the consumer use a conditional write with an idempotency key derived from the message ID.
C.Increase the table's provisioned write capacity units substantially and enable DynamoDB Streams so that duplicate items can be detected after the fact.
D.Move the table to a different AWS Region and enable DynamoDB Accelerator (DAX) in front of it to cache the writes during the spike.
AnswerB

On-demand capacity mode removes the need to forecast or preprovision read and write capacity, so the table scales with the traffic spike automatically. Using a conditional write keyed on a stable message identifier makes repeated processing of the same message a no-op, eliminating duplicates. Together these address both the throttling and the duplicate-order symptoms described.

Why this answer

The symptom set has two distinct causes: a capacity model that cannot absorb spikes and a consumer that is not idempotent. On-demand capacity mode lets the table follow the workload without manual provisioning, while a conditional write on a stable idempotency key makes repeated delivery harmless. Addressing both together is what stops the throttling and the duplicate orders.

Exam trap

The trap here is treating throttling and duplicate processing as one problem, when they require two independent fixes.

238
MCQeasy

Several EC2 instances in different Availability Zones need to read and write the same shared file system. The file storage should stay available if one AZ has a problem. Which service should the team choose?

A.Amazon EBS
B.Amazon EFS
C.Amazon S3 only
D.Instance store
AnswerB

Amazon EFS is a managed shared file system that can be mounted by multiple EC2 instances across multiple Availability Zones. It is a strong fit when applications need the same files at the same time and must remain available even if one AZ experiences issues. The service is highly available by design and reduces operational work compared with self-managed file servers.

Why this answer

Amazon EFS provides a fully managed, scalable, and elastic NFS file system that can be mounted concurrently by multiple EC2 instances across different Availability Zones. It is designed for high availability and durability by storing data redundantly across multiple AZs within a region, ensuring continued access even if one AZ fails.

Exam trap

The trap here is that candidates often confuse EBS Multi-Attach (which only supports a limited number of instances in the same AZ and requires a cluster-aware file system) with the true multi-AZ shared file system capability of EFS.

Why the other options are wrong

A

Amazon EBS volumes are tied to a single Availability Zone and cannot be shared across multiple EC2 instances in different AZs, so they cannot provide the required shared file system with cross-AZ availability.

C

Amazon S3 is object storage, not a shared file system; EC2 instances cannot mount S3 as a POSIX-compliant file system for concurrent read/write access. It also does not provide file locking or low-latency file operations needed for shared file systems.

When would these options actually be correct?

A

A single EC2 instance needs a durable, block-level storage volume with consistent low latency, and the workload does not require multi-instance access or cross-AZ resilience. For example, a database server's data volume.

C

When the requirement is to store and retrieve large amounts of unstructured data (e.g., images, videos, backups) from multiple EC2 instances, and the application uses S3 SDK or REST API rather than a mounted file system. Also correct for static website hosting or data lake scenarios.

Why candidates pick the wrong answer

A

Candidates may confuse EBS with a shared file system because EBS volumes can be attached to one instance at a time, but they overlook the multi-AZ and multi-instance sharing requirements.

C

Candidates may think S3 is highly available and durable, and assume it can serve as a shared file system because it is accessible from any instance, overlooking the lack of file system semantics like locking and POSIX compliance.

239
MCQmedium

A SaaS vendor will access your AWS resources by assuming an IAM role in your account. You want to prevent confused-deputy attacks and ensure the vendor can only assume the role using an agreed external identifier. Your role trust policy currently allows sts:AssumeRole from the vendor’s principal, but it does not include any external ID protection. Which change is the best next step?

A.Add a condition to the trust policy: Condition = {"StringEquals": {"sts:ExternalId": "vendor-agreed-id"}}.
B.Add a condition to the trust policy: Condition = {"IpAddress": {"aws:SourceIp": "203.0.113.0/24"}}.
C.Remove sts:AssumeRole and replace it with sts:AssumeRoleWithWebIdentity to use the vendor’s browser-based tokens.
D.Add a condition to the role permissions policy (not the trust policy) requiring aws:PrincipalTag/ExternalId to equal the external identifier.
AnswerA

Using sts:ExternalId in the trust policy ensures only assume-role requests presenting the correct external identifier are allowed. This directly mitigates confused-deputy attacks by binding authorization to a value the vendor must know. It also keeps the permissions model clean, because the check is enforced during the STS AssumeRole request.

Why this answer

The `sts:ExternalId` condition key is specifically designed to prevent confused-deputy problems. By adding `{"StringEquals": {"sts:ExternalId": "vendor-agreed-id"}}` to the trust policy, you ensure that the vendor must provide the agreed external ID in the `AssumeRole` API call, which only the legitimate vendor knows. This prevents a malicious third party from tricking the vendor into assuming a role in your account on their behalf.

Exam trap

The trap here is that candidates often confuse where to place the condition (trust policy vs. permissions policy) or mistakenly think IP-based restrictions or changing the API action are appropriate solutions for confused-deputy prevention.

Why the other options are wrong

B

The question requires protection against confused-deputy attacks using an external ID, not IP-based restrictions. The vendor's IP addresses may change or be shared, and IP conditions do not prevent a different vendor from using the same role.

C

This option is wrong because the question is about preventing confused-deputy attacks when a vendor assumes an IAM role, which requires sts:AssumeRole with an external ID condition, not sts:AssumeRoleWithWebIdentity, which is used for federated users with web identity tokens (e.g., from Amazon Cognito, Google, or Facebook).

D

The permissions policy controls what actions the role can perform, not who can assume it. The external ID check must be in the trust policy to prevent confused-deputy attacks during role assumption.

When would these options actually be correct?

B

This would be correct if the question asked to restrict role assumption to requests originating from a specific, static IP range owned by the vendor, such as when the vendor has a fixed corporate network and the goal is to limit access by network location.

C

This option would be correct in a scenario where a web application allows users to sign in via a third-party identity provider (e.g., Google or Facebook) and then accesses AWS resources using temporary credentials obtained through web identity federation. The trust policy would then use sts:AssumeRoleWithWebIdentity with conditions on the token's claims.

D

If the question asked how to restrict the role's actions based on a specific external identifier after assumption (e.g., logging or resource tagging), then a condition in the permissions policy using aws:PrincipalTag/ExternalId would be appropriate.

Why candidates pick the wrong answer

B

Candidates may think IP restriction is a general security best practice and assume it also prevents confused-deputy attacks, not realizing that external ID is the specific mechanism for that threat.

C

Candidates might confuse the need for an external identifier with web identity federation, thinking that using a web identity token provides a similar security mechanism, or they may not fully understand the difference between sts:AssumeRole and sts:AssumeRoleWithWebIdentity.

D

Candidates may confuse the purpose of trust policies vs. permissions policies, thinking that any condition related to the external ID can be placed in the permissions policy.

240
MCQhard

A warehouse integration service must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable?

A.Use CloudFront signed URLs
B.Use Amazon SQS standard queue and design consumers to be idempotent
C.Use UDP messages sent directly to workers
D.Use an in-memory queue on one EC2 instance
AnswerB

Amazon SQS standard queues provide at-least-once delivery with high throughput, meaning every message is delivered but occasional duplicates can occur. Designing consumers to be idempotent ensures that processing the same event multiple times yields the same result, which satisfies the requirement to process every event. This is the recommended pattern for reliable, scalable event processing in AWS.

Why this answer

Amazon SQS standard queues provide at-least-once delivery, meaning each message is delivered at least once but can occasionally be delivered more than once. This matches the requirement to process every event at least once, and since duplicate processing is acceptable when consumers are idempotent, the standard queue is the most suitable and cost-effective choice. SQS also decouples the warehouse integration service from its consumers, improving resilience and scalability.

Exam trap

The trap here is that candidates may confuse 'at-least-once' with 'exactly-once' and incorrectly choose FIFO queues or other options, but the question explicitly accepts duplicates if idempotency is handled, making the standard queue the correct and simpler choice.

How to eliminate wrong answers

Option A is wrong because CloudFront signed URLs are used to control access to content delivered via CloudFront, not for event processing or message queuing; they provide no delivery guarantee mechanism. Option C is wrong because UDP is a connectionless, unreliable transport protocol that does not guarantee message delivery, order, or duplicate prevention, making it unsuitable for at-least-once processing. Option D is wrong because an in-memory queue on a single EC2 instance creates a single point of failure and lacks durability; if the instance fails, all queued events are lost, violating the requirement to process every event at least once.

241
MCQhard

A company has a VPC with a CIDR block of 10.0.0.0/16. They need to allow an on-premises data center (192.168.0.0/24) to access a web application running on EC2 instances in a private subnet. The security team wants to ensure that only HTTP and HTTPS traffic from the on-premises network is allowed, and that the traffic is encrypted in transit. Which combination of AWS services should they use?

A.Set up an AWS Site-to-Site VPN connection and configure security groups to allow HTTP/HTTPS from 192.168.0.0/24.
B.Set up an AWS Transit Gateway with a VPN attachment and configure security groups to allow all traffic from 192.168.0.0/24.
C.Set up an AWS Client VPN endpoint and configure security groups to allow HTTP/HTTPS from the VPN client CIDR.
D.Set up an AWS Direct Connect connection and configure network ACLs to allow HTTP/HTTPS from 192.168.0.0/24.
AnswerA

An AWS Site-to-Site VPN provides an encrypted tunnel over the internet between the on-premises network and the VPC. Security groups on the EC2 instances can be configured to allow inbound HTTP (port 80) and HTTPS (port 443) only from the on-premises CIDR. This meets the requirements for encryption in transit and restricted traffic.

Why this answer

An AWS Site-to-Site VPN creates an encrypted tunnel between the on-premises network and the VPC, ensuring data in transit is protected. Security groups can then be configured to allow only HTTP and HTTPS traffic from the specific on-premises CIDR block, restricting access to the required ports. This combination satisfies both the encryption and traffic restriction requirements.

Exam trap

The trap here is assuming that AWS Direct Connect encrypts traffic by default; it does not, and would require an additional VPN for encryption.

242
MCQeasy

A team needs a relational database solution that can automatically fail over to a standby instance if the primary database becomes unavailable. They want the standby to be located in a different Availability Zone. Which RDS/Aurora configuration best satisfies this requirement?

A.Single-AZ DB deployment and rely on manual snapshot restore during failures.
B.Multi-AZ deployment with an automatically managed standby in a different Availability Zone and automatic failover.
C.Enable read replicas only, and promote a replica manually when the primary fails.
D.Enable point-in-time recovery (PITR) without configuring any Multi-AZ standby.
AnswerB

In a Multi-AZ deployment, Amazon RDS or Aurora provisions a physically separate standby instance in a different Availability Zone and synchronously replicates data from the primary to that standby. When the primary fails or the AZ is degraded, RDS automatically initiates failover to the standby, which becomes the new primary; because the DNS endpoint remains unchanged, the application continues to work without manual intervention. This provides both the required cross-AZ redundancy and automatic failover, meeting the availability requirement directly.

Why this answer

A Multi-AZ RDS deployment automatically provisions and maintains a standby instance in a different Availability Zone, and the failover is handled automatically by AWS without manual intervention. This meets the requirement for automatic failover to a standby in a different AZ, which is the core purpose of Multi-AZ deployments.

Exam trap

The trap here is that candidates often confuse read replicas with Multi-AZ standby, thinking that promoting a read replica provides automatic failover, but read replicas require manual promotion and do not serve as a synchronous standby.

How to eliminate wrong answers

Option A is wrong because a Single-AZ deployment has no standby instance, and manual snapshot restore requires significant downtime and manual steps, failing the automatic failover requirement. Option C is wrong because read replicas are designed for read scaling, not automatic failover; promoting a read replica manually introduces downtime and does not provide automatic failover to a standby. Option D is wrong because point-in-time recovery (PITR) only enables restoring to a specific time from backups, not automatic failover to a standby instance in a different AZ.

243
MCQeasy

Your web application runs on EC2 instances behind an Application Load Balancer (ALB). During traffic spikes, p95 response time increases, but average CPU utilization remains below 40%. The current Auto Scaling policy scales based on average CPU%. What should you change to improve performance during spikes?

A.Keep scaling on CPU% to avoid over-scaling
B.Scale on a request-driven metric such as ALB RequestCount per target (or target-group request rate)
C.Disable scaling and manually increase capacity during business hours
D.Scale only when network packet drops fall below a threshold
AnswerB

A request-driven metric correlates directly with incoming workload pressure. Scaling on request rate helps ensure enough capacity is added before request queues build up, which can reduce p95 response time even when CPU remains low.

Why this answer

The p95 response time is increasing during traffic spikes while CPU utilization remains low, indicating that the bottleneck is not compute capacity but rather request handling or connection overhead. By scaling on ALB RequestCountPerTarget, you directly target the metric causing latency—each target's request load—rather than an indirect metric like CPU. This ensures that new instances are launched precisely when individual targets are overwhelmed by requests, reducing queueing delays and improving response times.

Exam trap

The trap here is that candidates assume high latency always means high CPU, but AWS tests the understanding that p95 latency can spike due to request queueing even when CPU is idle, making request-based scaling the correct choice over CPU-based scaling.

How to eliminate wrong answers

Option A is wrong because continuing to scale on CPU% ignores the actual symptom (high p95 latency with low CPU), leading to under-provisioning during request bursts. Option C is wrong because manual scaling during business hours is not elastic and cannot react to unpredictable traffic spikes, violating the principle of auto scaling for performance. Option D is wrong because scaling on network packet drops is irrelevant to the described issue (low CPU, high latency) and packet drops typically indicate network congestion or buffer exhaustion, not request overload on the application layer.

244
MCQhard

A healthcare analytics platform processes streaming records with an AWS Lambda function that writes results to an Amazon DynamoDB table. The pipeline must not lose records if the function throws an error, and the operations team wants to inspect and reprocess failed records without writing custom retry code. Which approach should the solutions architect use?

A.Write a wrapper inside the function that catches exceptions and re-sends the batch to the stream before returning success
B.Configure the event source mapping with a maximum retry count and a destination on failure set to an Amazon SQS queue configured as a dead-letter queue
C.Set the function's timeout to the maximum value and rely on Lambda's built-in retry of the entire batch until it succeeds
D.Increase the Lambda function's reserved concurrency so that retries happen faster and failures are less likely
AnswerB

For Lambda event source mappings, the maximum retry count controls how many times a failing batch is retried, and the on-failure destination sends the batch metadata to an SQS queue or SNS topic after retries are exhausted. That preserves failed records for later inspection and reprocessing without custom code. This directly meets both the no-loss and no-custom-retry requirements.

Why this answer

Lambda event source mappings support a maximum retry count plus an on-failure destination that captures the failed batch after retries are exhausted. Pointing that destination at an SQS queue gives the team a durable holding area where failed records can be examined and replayed, satisfying both no-loss and no-custom-code goals. The other options either only tune performance, rely on nonexistent indefinite retries, or reintroduce the custom logic the team wants to eliminate.

Exam trap

The trap here is believing Lambda retries a failing batch indefinitely; event source mappings retry a limited number of times, and without an on-failure destination the batch is discarded once retries are exhausted.

245
MCQmedium

A financial services firm runs a critical API on Amazon EC2 instances behind a Network Load Balancer. The API must handle a sudden loss of one Availability Zone and continue serving traffic with no manual failover. The instances are in an Auto Scaling group that currently uses a single subnet in one Availability Zone. Which change should the architect make?

A.Replace the Network Load Balancer with an Application Load Balancer and enable sticky sessions with a long cookie duration.
B.Create an Amazon Route 53 latency-based routing record that points to the Network Load Balancer and set a failover routing policy with a health check.
C.Enable cross-zone load balancing on the Network Load Balancer and increase the Auto Scaling group desired capacity to four instances in the existing subnet.
D.Recreate the Auto Scaling group with subnets in at least two Availability Zones, enable the Network Load Balancer across those subnets, and attach a target group with health checks.
AnswerD

A Network Load Balancer requires subnets in each Availability Zone where it should accept traffic, and an Auto Scaling group spanning multiple AZs keeps instances running if one AZ fails. Health checks remove failed targets, so the API continues serving with no manual intervention, which is exactly the resilience the firm needs.

Why this answer

The Network Load Balancer must be enabled in subnets across multiple Availability Zones, and the Auto Scaling group must launch instances in those same subnets. With health checks on the target group, failed instances and the failed AZ are removed from rotation automatically, so the API remains available without manual failover.

Exam trap

The trap here is believing that cross-zone load balancing or a larger desired capacity creates Availability Zone redundancy, when both still depend on having subnets and instances in more than one AZ.

246
MCQmedium

A test environment runs on x86 EC2 instances and uses open-source software with no architecture-specific licensing restriction. What should be evaluated to reduce compute cost?

A.Cross-Region data replication for all data
B.AWS Graviton-based instances after performance testing
C.io2 Block Express volumes for all instances
D.Dedicated Hosts by default
AnswerB

AWS Graviton-based instances, such as m6g, c6g, and r6g, are built on the Arm architecture and typically offer up to 20% better price-performance than comparable x86 instances. The catch is that your application and its dependencies must be compiled or runnable on Arm, so performance testing is essential to confirm compatibility and measure actual throughput, latency, and cost-per-request before migrating the test environment.

Why this answer

AWS Graviton-based instances (e.g., M6g, C6g) use Arm-based custom AWS silicon, offering up to 40% better price-performance compared to comparable x86 instances for many workloads. Since the test environment runs open-source software with no architecture-specific licensing restrictions, migrating to Graviton after performance testing can significantly reduce compute costs without compatibility issues.

Exam trap

The trap here is that candidates may assume all cost optimization involves reducing instance size or using Spot Instances, but the question specifically tests knowledge of architecture-specific cost savings with Graviton when no licensing restrictions exist.

How to eliminate wrong answers

Option A is wrong because cross-Region data replication increases data transfer and storage costs, not compute costs, and is a data durability/disaster recovery feature, not a cost optimization for compute. Option C is wrong because io2 Block Express volumes are high-performance, high-cost SSD volumes designed for latency-sensitive workloads, not a compute cost reduction strategy; they would increase storage costs unnecessarily for a test environment. Option D is wrong because Dedicated Hosts incur additional per-host charges and are intended for licensing or compliance requirements (e.g., Windows Server with dedicated licensing), not for general compute cost reduction; using them by default would increase costs.

247
MCQmedium

A trading dashboard runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The architecture review board prefers a managed AWS-native control.

A.A single EC2 instance with detailed monitoring
B.Subnets in at least two Availability Zones with health checks enabled
C.All instances in one larger subnet
D.A Network Load Balancer in one subnet
AnswerB

Placing the Auto Scaling group's subnets in at least two Availability Zones ensures that if one AZ becomes unavailable, the remaining AZs still have healthy instances to serve traffic. Health checks (either EC2 status checks or Elastic Load Balancing target health checks) allow the ASG to detect failed instances and replace them while maintaining the desired capacity. This architecture survives both individual instance crashes and full AZ outages, providing high availability for the trading dashboard.

Why this answer

Distributing EC2 instances across at least two Availability Zones (AZs) ensures that the application remains available if one AZ fails. The Auto Scaling group must include subnets in multiple AZs and use health checks (e.g., ELB health checks) to automatically replace unhealthy instances. This configuration meets the requirement for fault tolerance and aligns with AWS-managed best practices for high availability.

Exam trap

The trap here is that candidates often confuse 'scaling' with 'resilience' and think that a single large subnet or a different load balancer type (NLB) provides AZ fault tolerance, but only multi-AZ subnet configuration with health checks ensures automatic recovery from an AZ failure.

How to eliminate wrong answers

Option A is wrong because a single EC2 instance, even with detailed monitoring, cannot tolerate the failure of an Availability Zone; it represents a single point of failure. Option C is wrong because placing all instances in one larger subnet confines them to a single Availability Zone, which does not provide AZ-level fault tolerance. Option D is wrong because a Network Load Balancer in one subnet does not address the need for multi-AZ instance distribution; it also lacks the health-check-based auto-scaling capabilities required for instance replacement.

248
MCQmedium

A retail company runs an e-commerce platform on a fleet of Amazon EC2 instances behind an Application Load Balancer. Traffic follows a predictable pattern: high during business hours and very low overnight. The operations team wants to reduce EC2 costs without affecting availability during peak hours. The instances currently run continuously and are managed by an Auto Scaling group with a minimum capacity of 4 and a maximum of 20. Which solution will meet these requirements MOST cost-effectively?

A.Enable detailed monitoring on all instances and create a target tracking scaling policy based on CPU utilization.
B.Purchase 4 Standard Reserved Instances for a 3-year term and let the Auto Scaling group launch On-Demand instances beyond that baseline.
C.Replace the Auto Scaling group with a larger number of smaller instances to improve granularity of scaling.
D.Configure a scheduled scaling action on the Auto Scaling group to reduce the desired capacity during off-peak hours and increase it before peak hours.
AnswerD

Scheduled scaling lets you set the desired capacity based on a known schedule, so the Auto Scaling group can scale down when traffic is predictably low and scale up before the peak. This directly matches capacity to demand and avoids paying for idle instances during off-peak hours, making it the most cost-effective option for a predictable pattern.

Why this answer

For workloads with a predictable daily or weekly traffic pattern, scheduled scaling is the most direct way to align capacity with demand. It reduces the number of running instances during known low-traffic periods without waiting for a metric-based policy to react. Reserved Instances or target tracking do not eliminate the cost of idle overnight capacity in this scenario.

Exam trap

The trap here is assuming that a discount purchasing option or a reactive scaling policy will automatically reduce cost, when the real issue is that the fleet keeps running at full baseline capacity during predictably low-traffic hours.

249
MCQhard

A media archive needs low-latency full-text search across product descriptions and filtered attributes. Which managed service is most suitable? The design must avoid adding custom operational scripts.

A.AWS Config
B.Amazon OpenSearch Service
C.Amazon EFS
D.Amazon SQS
AnswerB

Amazon OpenSearch Service provides managed full-text search with low-latency querying across document fields, plus filter clauses for structured attributes. It satisfies the no-custom-operational-scripts constraint because AWS handles cluster provisioning, patching and scaling, unlike self-managed search engines requiring bespoke maintenance automation.

Why this answer

Amazon OpenSearch Service is the correct choice because it provides managed, low-latency full-text search capabilities with support for filtering on structured attributes (e.g., product categories, price ranges). It indexes JSON documents and exposes a RESTful API for search queries, eliminating the need for custom operational scripts while meeting the media archive's requirements.

Exam trap

The trap here is that candidates might confuse AWS Config's resource tracking or EFS's file storage with search capabilities, overlooking that OpenSearch Service is the only managed option purpose-built for full-text search and filtering.

How to eliminate wrong answers

Option A is wrong because AWS Config is a service for auditing and evaluating resource configurations against compliance rules, not for full-text search or indexing product descriptions. Option C is wrong because Amazon EFS is a scalable NFS file system for shared storage, not a search engine; it cannot perform low-latency full-text queries across text content. Option D is wrong because Amazon SQS is a managed message queue for decoupling application components, not a search or indexing service, and it does not support querying stored data.

250
MCQmedium

An event-driven order processing service consumes messages from an Amazon SQS Standard queue. After a deployment, about 1% of messages start failing validation because a required field is missing. The consumer catches the exception and returns control, so the messages are retried. However, those poison messages keep reappearing and repeatedly consuming processing time for hours, delaying handling of valid messages. What is the most resilient way to handle the poison messages while keeping the system available?

A.Set the consumer visibility timeout to a very large value so failing messages are hidden for hours.
B.Configure an SQS redrive policy to send messages to a dead-letter queue (DLQ) after a limited number of receives (maxReceiveCount).
C.Switch the SQS queue from Standard to FIFO so poison messages do not retry.
D.Increase the consumer concurrency indefinitely so the system processes all messages even if some fail validation.
AnswerB

A DLQ redrive policy creates a deterministic stop condition for poison messages. After maxReceiveCount, the messages are moved to the DLQ instead of cycling in the main queue, preventing repeated failed deliveries from degrading capacity and availability for valid messages.

Why this answer

Configuring an SQS redrive policy with a maxReceiveCount (e.g., 3–5) automatically moves messages that repeatedly fail processing to a dead-letter queue (DLQ) after the specified number of receives. This isolates the poison messages, preventing them from consuming visibility timeout and processing resources, while allowing valid messages to be handled without delay. The DLQ can then be analyzed or reprocessed offline, maintaining system availability.

Exam trap

The trap here is that candidates may think increasing visibility timeout or concurrency solves the problem, but they fail to recognize that only a dead-letter queue permanently isolates poison messages from the processing pipeline.

How to eliminate wrong answers

Option A is wrong because setting the consumer visibility timeout to a very large value would hide failing messages for hours, but they would still reappear after the timeout expires, continuing the cycle of retries and delays without resolving the issue. Option C is wrong because switching from Standard to FIFO does not prevent poison messages from retrying; FIFO queues still retry messages on failure and require a DLQ for poison handling, and they also sacrifice throughput and ordering flexibility. Option D is wrong because increasing consumer concurrency indefinitely does not address the root cause—poison messages will still be retried and consume processing slots, potentially overwhelming the system and delaying valid messages further.

251
MCQhard

A order processing API must ensure that only encrypted EBS volumes can be created in the account. What is the strongest preventive control?

A.Run a daily Lambda function to encrypt unencrypted volumes
B.Enable VPC Flow Logs
C.Use an SCP that denies ec2:CreateVolume when the encrypted condition is false
D.Tag encrypted volumes after creation
AnswerC

An SCP denying `ec2:CreateVolume` when the `ec2:Encrypted` condition key is false blocks unencrypted volume creation across every account in the organisation, regardless of IAM permissions. This preventive guardrail satisfies the requirement that only encrypted EBS volumes can be created, unlike detective controls or default encryption settings.

Why this answer

Service Control Policies (SCPs) are a preventive control that can deny the ec2:CreateVolume API call when the encryption condition (ec2:Encrypted) is false. This ensures that no unencrypted EBS volumes can be created at the account level, regardless of IAM permissions. SCPs operate at the AWS Organizations root, OU, or account level and are evaluated before any IAM policies, making them the strongest preventive mechanism.

Exam trap

The trap here is confusing detective/reactive controls (like Lambda remediation) with preventive controls (like SCPs), leading candidates to choose a solution that fixes the problem after it occurs rather than blocking it entirely.

How to eliminate wrong answers

Option A is wrong because running a daily Lambda function to encrypt unencrypted volumes is a detective/reactive control, not a preventive one; it does not block the creation of unencrypted volumes and leaves a window of exposure. Option B is wrong because VPC Flow Logs capture network traffic metadata (IP addresses, ports, protocols) and have no ability to enforce encryption policies on EBS volumes; they are a monitoring tool, not a preventive control. Option D is wrong because tagging encrypted volumes after creation is a labeling action that does not prevent unencrypted volumes from being created; it is a detective or organizational control, not a preventive one.

252
MCQhard

A claims workflow uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The architecture review board prefers a managed AWS-native control.

A.A FIFO queue without a redrive policy
B.Short polling instead of long polling
C.A dead-letter queue with an appropriate maxReceiveCount
D.A larger message retention period only
AnswerC

A dead-letter queue with maxReceiveCount moves messages aside once they exceed the receive threshold, so poison messages stop blocking useful retries. This satisfies the review board's managed AWS-native constraint, since DLQs are a native SQS feature requiring no custom code or third-party tooling.

Why this answer

A dead-letter queue (DLQ) with an appropriate maxReceiveCount allows messages that repeatedly fail processing to be moved out of the source queue after a specified number of receive attempts. This prevents poison messages from blocking useful retries and is a fully managed AWS-native pattern. The architecture review board's preference for a managed solution is satisfied because SQS DLQs are a built-in feature requiring no custom code.

Exam trap

The trap here is that candidates may confuse a DLQ with simply increasing retention or changing polling behavior, not realizing that poison messages require explicit isolation via a separate queue and a maxReceiveCount threshold to stop infinite retries.

How to eliminate wrong answers

Option A is wrong because a FIFO queue without a redrive policy does not automatically handle poison messages; without a DLQ, failed messages remain in the queue and continue to block retries. Option B is wrong because short polling reduces latency but does not address poison messages; it returns only a subset of servers' messages and can increase empty responses, but it has no effect on message failure handling. Option D is wrong because increasing the message retention period only keeps messages longer without removing failing ones; poison messages would still be retried until they expire, continuing to block useful retries.

253
Multi-Selectmedium

An application uses an Amazon RDS Multi-AZ DB instance. During a failover test, connections fail until the application is restarted, even though the database comes back online. Which two changes should the team make to improve resilience during failover? Select two.

Select 2 answers
A.Cache and reconnect to the current writer IP address to avoid DNS lookups during failover.
B.Use the RDS endpoint name instead of hard-coding the current instance IP or hostname in the application.
C.Switch to a read replica and let it promote manually after every outage.
D.Add retry logic with exponential backoff for transient connection and DNS resolution errors.
E.Disable connection pooling so each request opens a fresh socket during normal operation.
AnswersB, D

The RDS endpoint abstracts the underlying writer instance. When failover occurs, AWS updates the endpoint to point at the new writer, so the application should reconnect by using the managed name rather than a fixed IP or hostname.

Why this answer

The RDS endpoint is a DNS name that automatically resolves to the current writer instance's IP address. During a failover, the DNS record is updated to point to the new primary, so using the endpoint instead of a hard-coded IP or hostname allows the application to reconnect without manual intervention. Option D is correct because adding retry logic with exponential backoff handles transient failures during DNS resolution and connection establishment, which are common during the brief period when the DNS TTL has not yet expired after a failover.

Exam trap

The trap here is that candidates often think caching the IP (Option A) improves performance, but it actually breaks failover resilience because the application never learns the new writer's address after a failover.

254
MCQeasy

A startup runs a public-facing web application on Amazon EC2 instances behind an Application Load Balancer. The security team wants to protect the application from common web exploits such as SQL injection and cross-site scripting, and also wants to rate-limit requests from specific IP addresses. Which AWS service should be used to meet these requirements?

A.Amazon GuardDuty
B.AWS Shield Advanced
C.AWS WAF
D.AWS Network Firewall
AnswerC

AWS WAF inspects HTTP and HTTPS requests and can block common exploits such as SQL injection and cross-site scripting using managed rule groups. It also supports rate-based rules that count requests from a source IP over a time window, which meets the rate-limiting requirement. Associating a web ACL with the Application Load Balancer provides the needed protection.

Why this answer

AWS WAF is the service designed to filter and monitor HTTP requests at the application layer. It provides managed rule groups that block SQL injection and cross-site scripting, and rate-based rules that limit requests from specific IP addresses. Associating a web ACL with the Application Load Balancer enforces these protections directly on incoming traffic, meeting both requirements.

Exam trap

The trap here is confusing DDoS protection with application-layer exploit protection, when AWS Shield Advanced addresses volumetric attacks while AWS WAF handles HTTP-level filtering and rate limiting.

255
Multi-Selectmedium

A data lake stores raw files in a single Amazon S3 bucket that is shared by three internal analytics teams. Each team should access only its own prefix, and the company wants to eliminate ACL management because objects come from multiple producers. Which three changes should the architect make? Select three.

Select 3 answers
A.Create a separate S3 access point for each team and scope it to that team’s prefix.
B.Leave ACLs enabled so each producer can grant permissions directly on uploaded objects.
C.Set Object Ownership to Bucket owner enforced so ACLs are disabled.
D.Use bucket or access point policies to restrict access to the allowed principals and prefixes.
E.Make the bucket public and rely on application-layer authorization for data protection.
AnswersA, C, D

An S3 access point gives each team a distinct hostname and policy scoped to its own prefix, so a single shared bucket can be partitioned without duplicating data. This enforces per-team prefix isolation while access point policies replace per-object ACL grants.

Why this answer

Option A is correct because S3 access points provide a dedicated endpoint per team, and each access point can be scoped with a policy limited to that team's prefix, giving clean per-team isolation without duplicating buckets. Option C is correct because setting Object Ownership to Bucket owner enforced disables ACLs entirely, so the bucket owner automatically owns every object and ACL management is eliminated, which matches the requirement that objects come from multiple producers. Option D is correct because bucket policies and access point policies are the IAM-based mechanism that restricts each team to its allowed principals and prefixes once ACLs are disabled.

Option B is wrong because leaving ACLs enabled keeps the ACL management burden the company wants to remove. Option E is wrong because making the bucket public exposes the data and application-layer authorization does not replace S3-level access control.

Exam trap

The trap here is that candidates may think ACLs are necessary for multi-producer environments, but AWS recommends disabling ACLs and using bucket policies or access point policies with Object Ownership set to 'Bucket owner enforced' to simplify access control.

Why the other options are wrong

B

Leaving ACLs enabled contradicts the requirement to eliminate ACL management, and ACLs do not restrict access by prefix—they grant permissions on individual objects, which is not scalable for multiple producers and teams.

When would these options actually be correct?

B

In a scenario where objects are uploaded by a single producer and each object needs individual permissions (e.g., a shared bucket with per-object access control for different users), and the company is willing to manage ACLs.

Why candidates pick the wrong answer

B

Candidates may think ACLs provide a straightforward way for producers to control access to their uploaded objects, overlooking the management overhead and the requirement to avoid ACLs.

256
Multi-Selecthard

A solutions architect is designing a high-performance computing (HPC) workload that requires a shared file system with high throughput and low latency for thousands of compute instances. The workload also requires a caching layer to accelerate repeated reads of the same data. Which two AWS services should be combined to meet these requirements? (Choose two.)

Select 2 answers
A.Amazon FSx for Lustre
B.Amazon ElastiCache for Redis
C.Amazon S3 Glacier
D.Amazon EBS Multi-Attach enabled io1 volumes
E.AWS Storage Gateway File Gateway
AnswersA, B

Amazon FSx for Lustre is a high-performance file system optimized for HPC workloads. It provides sub-millisecond latencies, millions of IOPS, and hundreds of gigabytes per second of throughput. It integrates natively with Amazon S3, allowing data to be lazily loaded from S3 and written back. FSx for Lustre is designed for compute-intensive workloads and can be linked to an S3 bucket as a data repository. This makes it ideal for the shared file system requirement.

Why this answer

Amazon FSx for Lustre provides the high-performance shared file system needed for HPC, with low latency and high throughput. Amazon ElastiCache for Redis adds an in-memory caching layer to accelerate repeated reads, reducing load on the file system. Together, they meet the requirements for a scalable, high-performance HPC storage and caching solution.

Other options either have high latency (Glacier), are for hybrid scenarios (File Gateway), or lack the necessary scalability (EBS Multi-Attach).

Exam trap

The trap here is assuming that any shared storage service can handle HPC scale; services like EBS Multi-Attach or Storage Gateway are not designed for thousands of instances or low-latency HPC.

257
Multi-Selecthard

A latency-sensitive video platform uploads large files to S3 from users around the world. Which two features can improve upload performance? The architecture review board prefers a managed AWS-native control.

Select 2 answers
A.S3 Object Lock
B.S3 Transfer Acceleration
C.S3 multipart upload
D.S3 Inventory
AnswersB, C

Transfer Acceleration routes uploads through AWS edge locations to the nearest S3 endpoint over AWS's optimised backbone, cutting latency for globally distributed users. It is fully AWS-managed, meeting the review board's AWS-native constraint while improving worldwide upload performance.

Why this answer

S3 Transfer Acceleration (B) uses AWS edge locations to accelerate uploads over long distances by routing traffic through the AWS global network, reducing latency and packet loss compared to the public internet. Multipart upload (C) improves performance by splitting large files into smaller parts that can be uploaded in parallel, increasing throughput and allowing retries of individual parts without restarting the entire upload.

Exam trap

The trap here is that candidates may confuse S3 Transfer Acceleration with CloudFront or think multipart upload is only for reliability, not performance, while overlooking that both features are managed AWS-native controls that directly address latency and throughput for large file uploads.

258
MCQmedium

An Auto Scaling group behind an Application Load Balancer frequently replaces new EC2 instances. The application needs ~6 minutes to warm up after instance launch. However, the ALB target group health checks start immediately and mark the targets unhealthy until the application is ready. Because the targets become unhealthy early, the Auto Scaling group then terminates the instances and launches replacements, creating a repeated unhealthy/termination loop. What configuration change will most directly improve recovery by preventing premature ASG termination while the application is warming up?

A.Set a health check grace period on the Auto Scaling group that exceeds the application startup/warm-up time.
B.Increase the Auto Scaling group's desired capacity to a higher number than required.
C.Disable ALB target group health checks so instances are considered healthy as soon as they register.
D.Change the Auto Scaling health check type from ELB to EC2 so the ALB will no longer determine instance health.
AnswerA

A health check grace period delays when the Auto Scaling group starts evaluating instance health. This prevents the ASG from terminating instances due to ALB/target health being unhealthy during the initial warm-up window, breaking the unhealthy/termination loop.

Why this answer

The health check grace period on an Auto Scaling group (ASG) allows a newly launched EC2 instance to bypass health check failures for a specified duration. By setting this grace period to exceed the application's ~6-minute warm-up time, the ASG will not prematurely terminate the instance based on ALB health check results. This directly breaks the unhealthy/termination loop while the application initializes.

Exam trap

The trap here is that candidates may think disabling health checks or changing the health check type is a valid fix, but the correct solution is to use the ASG's built-in grace period to decouple early health check failures from termination decisions.

Why the other options are wrong

B

Increasing desired capacity does not prevent the Auto Scaling group from terminating instances that fail health checks; it only adds more instances, which may also fail and be terminated, perpetuating the loop.

C

Disabling ALB target group health checks would prevent the ALB from routing traffic to healthy instances, causing service disruption. The issue is premature termination by ASG, not health check failure; the grace period directly addresses this.

D

Changing the health check type to EC2 would make the Auto Scaling group ignore ALB health check results, but the ALB would still route traffic to unhealthy instances, causing application errors. The question requires preventing premature termination during warm-up, not ignoring health checks entirely.

When would these options actually be correct?

B

This option would be correct in a scenario where the application requires a minimum number of healthy instances to handle traffic, and the current desired capacity is too low to meet demand, causing performance issues or scaling events.

C

If the application has no external dependencies and health checks are causing false negatives due to a bug or misconfiguration, disabling them temporarily for troubleshooting or during a migration where health checks are not yet reliable could be correct.

D

This option would be correct if the question stated that the ALB health checks are misconfigured (e.g., checking a port that is not open) and the instances are actually healthy, but the ALB incorrectly marks them unhealthy. In that case, switching to EC2 health checks would prevent unnecessary terminations.

Why candidates pick the wrong answer

B

Candidates might think that adding more instances will compensate for the ones being terminated, but this ignores the root cause of premature termination due to health checks.

C

Candidates may think that eliminating health checks stops the termination loop, but they overlook that health checks are essential for traffic routing and that the grace period is the designed solution for warm-up delays.

D

Candidates may think that bypassing ALB health checks will stop the termination loop, but they overlook that the ALB still needs to know instance health for routing traffic, and the real issue is the warm-up time, not the health check type.

259
MCQeasy

A startup runs a stateless web application on Amazon EC2 instances behind an Application Load Balancer. Traffic is steady during the day but drops to almost zero overnight, and the team wants to reduce compute cost without manual intervention or a service interruption. Which action should the team take?

A.Replace the Auto Scaling group with a single large EC2 instance and use Elastic Load Balancing health checks.
B.Create a scheduled scaling action on an Auto Scaling group that reduces the desired capacity at night and restores it each morning.
C.Attach an Amazon EBS volume with higher IOPS to each instance to reduce the number of instances needed.
D.Switch to Spot Instances for the entire fleet and rely on capacity-optimized allocation.
AnswerB

Scheduled scaling adjusts the desired capacity of an Auto Scaling group at defined times, so the fleet can shrink overnight when traffic is near zero and grow before the morning peak. Because the group manages instance lifecycle and the ALB drains connections, there is no service interruption. This matches a predictable daily pattern and removes manual intervention.

Why this answer

The workload has a predictable daily pattern, which is exactly what scheduled scaling on an Auto Scaling group is designed for. Reducing desired capacity overnight and restoring it in the morning lowers compute spend while the group and load balancer keep the application available. Connection draining on the ALB ensures in-flight requests finish before instances terminate, so users see no interruption.

Exam trap

The trap here is choosing Spot Instances purely for cost when the scenario's real requirement is matching capacity to a known daily traffic curve without risking reclamation.

260
MCQeasy

A compute workload uses temporary scratch space for intermediate results (reproducible), and it can tolerate data loss if the instance is terminated. The workload benefits from very high local I/O throughput. Which storage option is the best fit for the scratch data?

A.Amazon EBS General Purpose (gp3) volumes to persist intermediate results across reboots.
B.Amazon EFS for a shared file system between multiple instances.
C.Instance store for local temporary files that can be lost when the instance stops.
D.Amazon S3 for scratch data so it is always durable and accessible from anywhere.
AnswerC

Instance store provides physically attached NVMe SSDs delivering the highest local I/O throughput and lowest latency, unlike EBS network storage. Because the intermediate results are reproducible and loss on termination is acceptable, the ephemeral, non-persistent nature of instance store matches the workload's tolerance exactly.

Why this answer

Instance store volumes provide very high local I/O throughput because they are physically attached to the host server, making them ideal for temporary scratch data that is reproducible and can tolerate loss. Since the workload explicitly accepts data loss on instance termination and does not require persistence across reboots, instance store is the best fit for this use case.

Exam trap

The trap here is that candidates often choose EBS gp3 (Option A) because they assume all block storage is persistent and high-performance, overlooking the fact that instance store offers even higher local throughput and is explicitly designed for temporary, loss-tolerant workloads.

Why the other options are wrong

B

Amazon EFS provides a shared file system, but the question specifies scratch data for a single instance that benefits from very high local I/O throughput. EFS is network-attached and has higher latency than local storage, making it unsuitable for high local I/O needs.

D

Amazon S3 is designed for durable, highly available object storage with high latency, not for high local I/O throughput scratch space. It cannot provide the very high local I/O performance required for temporary scratch data.

When would these options actually be correct?

B

A question where multiple instances need to concurrently access and share temporary files with low administrative overhead, and the workload can tolerate network latency. For example, a distributed data processing job that requires a common scratch space across nodes.

D

A scenario where the workload requires durable, scalable, and accessible storage for data that must persist across instance terminations and be shared across multiple applications or regions, such as storing backup files or static website assets.

Why candidates pick the wrong answer

B

Candidates may think a shared file system is beneficial for any temporary data, overlooking that the question emphasizes local I/O throughput and single-instance scratch space, not multi-instance sharing.

D

Candidates may think S3's durability and accessibility make it suitable for any data, overlooking the specific need for high local I/O throughput and the tolerance for data loss in this scratch data use case.

261
MCQmedium

A company runs an internet-facing API in two AWS Regions. Route 53 currently uses simple routing to a primary Application Load Balancer (ALB) DNS name. When the primary Region experiences an outage, customers wait a long time because the DNS entry is not changed automatically. The team wants automatic failover: if the primary Region ALB health check fails for a sustained period, Route 53 should route users to the secondary Region ALB. Which Route 53 approach best meets this requirement?

A.Use Route 53 failover routing with a PRIMARY and SECONDARY record set for the same name, and attach health checks to the ALBs.
B.Use latency-based routing so Route 53 automatically spreads traffic to both Regions based on measured latency.
C.Use weighted routing and configure the secondary ALB to receive 100% traffic when the primary returns HTTP 5xx responses.
D.Use geolocation routing and restrict the primary Region record to specific countries only.
AnswerA

Route 53 failover routing is specifically designed for active-passive DNS failover. You create two records with the same name, designate one as PRIMARY and one as SECONDARY, and attach a Route 53 health check to each ALB endpoint. Route 53 continuously evaluates the PRIMARY health check; when it fails for the configured evaluation period, Route 53 responds with the SECONDARY record's IP or alias. This gives a deterministic, health-driven failover where the healthy secondary ALB starts receiving traffic once the primary is marked unhealthy, respecting the record TTL for propagation.

Why this answer

Route 53 failover routing is designed specifically for active-passive failover scenarios. By creating PRIMARY and SECONDARY record sets with the same DNS name and attaching health checks to the ALBs, Route 53 will automatically route traffic to the secondary ALB when the primary ALB health check fails for a sustained period. This meets the requirement for automatic failover without manual intervention.

Exam trap

The trap here is that candidates often confuse failover routing with latency-based or weighted routing, assuming that latency-based routing inherently provides failover, but it does not—it only optimizes for performance, not availability.

Why the other options are wrong

B

Latency-based routing distributes traffic based on lowest latency, not health. It does not provide automatic failover when a region is completely down; users may still be routed to the unhealthy primary if it has lower latency.

C

Weighted routing distributes traffic based on weights, not health. It cannot automatically shift 100% traffic to the secondary ALB based on HTTP 5xx responses; health checks are not integrated with weighted routing for automatic failover.

D

Geolocation routing directs traffic based on the geographic location of the user, not on the health or availability of the endpoint. It cannot automatically failover to a secondary Region when the primary ALB becomes unhealthy.

When would these options actually be correct?

B

A company wants to route users to the region with the lowest latency for better performance, and both regions are healthy and active. They do not require failover; they simply want to optimize response times.

C

When you need to gradually shift traffic from one endpoint to another (e.g., blue/green deployment) or split traffic across multiple endpoints for A/B testing, and you manually adjust weights. Health checks are not required for this scenario.

D

A company needs to restrict access to its API based on the user's country due to licensing or regulatory requirements. For example, users from the EU must be routed to a specific ALB in Frankfurt, while users from the US go to an ALB in Virginia.

Why candidates pick the wrong answer

B

Candidates may think latency routing inherently handles failover because it 'automatically' routes to the best region, but it lacks health check awareness and can still direct traffic to an unhealthy endpoint.

C

Candidates may think weighted routing can be used for failover by setting weights to 0 and 100, but they overlook that Route 53 does not automatically adjust weights based on endpoint health.

D

Candidates may confuse geolocation routing with failover routing, thinking that restricting traffic to specific countries can somehow trigger a failover, or they may overestimate Route 53's ability to automatically detect and react to regional outages with geolocation policies.

262
MCQmedium

A global mobile game backend serves mostly static images and JavaScript files from an S3 origin. Users in distant countries report slow load times. What should improve performance most? The architecture review board prefers a managed AWS-native control.

A.RDS read replicas
B.Amazon CloudFront distribution with the S3 bucket as origin
C.A larger S3 bucket
D.An EC2 Auto Scaling group in one Region
AnswerB

CloudFront caches the static images and JavaScript at edge locations close to distant users, cutting latency versus direct S3 retrieval. It is fully managed and AWS-native, satisfying the review board's preference, and integrates directly with the S3 bucket as origin.

Why this answer

Amazon CloudFront is a global content delivery network (CDN) that caches static content (images, JavaScript) at edge locations close to users, drastically reducing latency. By using the S3 bucket as the origin, CloudFront offloads requests from S3 and serves cached objects from the nearest edge, which directly addresses slow load times for distant users. This is a managed AWS-native service that aligns with the architecture review board's preference.

Exam trap

The trap here is that candidates may think increasing S3 bucket size or using RDS replicas can improve static content delivery, but the core issue is geographic latency, which only a CDN like CloudFront can solve by caching content at edge locations.

How to eliminate wrong answers

Option A is wrong because RDS read replicas are designed to offload read traffic from a relational database, not to accelerate delivery of static files stored in S3; they have no effect on S3 latency. Option C is wrong because increasing the S3 bucket size does not improve data transfer speed or reduce latency; S3 performance is independent of bucket size and is limited by regional endpoints. Option D is wrong because an EC2 Auto Scaling group in a single Region does not provide geographic distribution; users in distant countries would still experience high latency connecting to that single Region, and it adds unnecessary compute overhead for serving static content.

263
MCQmedium

You use Amazon CloudFront in front of a private content S3 origin. To mitigate an OWASP Top 10 issue, you created a WAF web ACL and associated it to the CloudFront distribution, but attacks are still reaching the origin. CloudWatch logs show the web ACL rules never match for the CloudFront requests. What is the most likely configuration mistake?

A.The WAF web ACL intended for CloudFront must be created in the us-east-1 (N. Virginia) region (CloudFront scope), even if the rest of the stack is in another region.
B.WAF rules only evaluate requests after they reach the origin, so the absence of matches means the origin is blocking traffic first.
C.For CloudFront, you must use a regional WAF endpoint and cannot use a global web ACL.
D.WAF web ACL rules never apply to signed URLs or signed cookies, so the web ACL is bypassed by design.
AnswerA

CloudFront-scoped WAF web ACLs use a global scope that is provisioned/managed in us-east-1. Creating the web ACL in the wrong region (or with the wrong scope) prevents CloudFront from evaluating the expected web ACL rules, which would lead to no rule matches in logs.

Why this answer

When using AWS WAF with CloudFront, the web ACL must be created in the US East (N. Virginia) region (us-east-1) because CloudFront is a global service that only supports WAF web ACLs with a global scope, which are always defined in us-east-1. If the web ACL is created in any other region, it will be a regional web ACL and cannot be associated with a CloudFront distribution, causing the rules to never be evaluated against incoming requests.

This explains why CloudWatch logs show no rule matches—the web ACL is effectively not attached to the CloudFront distribution.

Exam trap

The trap here is that candidates assume WAF web ACLs can be created in any region for CloudFront, not realizing that CloudFront requires a global-scope web ACL that must be created in us-east-1, regardless of where the origin or other resources reside.

Why the other options are wrong

B

WAF rules evaluate requests before they reach the origin, not after. The absence of matches indicates the web ACL is not being applied to CloudFront traffic, not that the origin is blocking requests.

C

CloudFront requires a global (CloudFront scope) web ACL, not a regional one. Associating a regional WAF web ACL with CloudFront is not supported, but the mistake here is that the web ACL was created in the wrong region (not us-east-1), not that it was regional.

D

WAF rules do apply to requests using signed URLs or signed cookies; the web ACL evaluates all requests that reach CloudFront, regardless of authentication method. The issue here is that the web ACL is not being applied at all because it was created in the wrong region.

When would these options actually be correct?

B

If the question described a scenario where the origin (e.g., an ALB) has its own WAF web ACL that is blocking traffic before CloudFront's WAF rules are evaluated, then option B could be correct in the sense that the origin's WAF is blocking requests first, but the statement as written is still inaccurate because WAF rules evaluate before reaching the origin.

C

If the question stated that you are using an Application Load Balancer (ALB) or API Gateway in a specific region, then you must create a regional WAF web ACL in that same region and associate it with the ALB or API Gateway.

D

If a question states that a WAF web ACL is associated with a CloudFront distribution but requests with signed URLs are still bypassing the WAF and reaching the origin, and the WAF logs show no matches, then the correct answer could be that signed URLs bypass WAF evaluation. However, this is incorrect in practice; WAF evaluates all requests. A more plausible scenario: a question about CloudFront signed URLs and WAF might incorrectly claim that WAF does not inspect signed URLs, but that would be a trick.

Why candidates pick the wrong answer

B

Candidates may confuse the order of evaluation or think that WAF only inspects traffic at the origin, not realizing that CloudFront integrates with WAF to inspect at the edge before forwarding to the origin.

C

Candidates may confuse the regional vs. global scope of WAF, thinking CloudFront can use a regional endpoint, or they may not know that CloudFront requires a global web ACL created in us-east-1.

D

Candidates may confuse signed URLs/cookies with authentication mechanisms that bypass WAF, thinking that pre-signed URLs skip WAF inspection, when in fact WAF evaluates all requests at the CloudFront edge.

264
MCQmedium

A logistics company runs a REST API on Amazon ECS using the Fargate launch type behind an Application Load Balancer. The API's response times are acceptable, but the operations team wants to reduce the number of database calls per request by caching frequently accessed reference data in memory inside the tasks. The data changes infrequently and slight staleness is acceptable. Which approach best meets these requirements?

A.Store the reference data in an Amazon ElastiCache for Redis cluster and have tasks query it on each request.
B.Implement an in-process cache in the application code that loads reference data at task startup and refreshes it periodically.
C.Enable Amazon ElastiCache for Memcached and configure automatic discovery of cache nodes.
D.Use the ECS task definition to mount an Amazon EFS file system containing the reference data.
AnswerB

An in-process cache keeps the reference data in the task's own memory, so requests are served without any database call or network hop. Since the data changes infrequently and slight staleness is acceptable, a periodic refresh is sufficient and simple. This directly satisfies the goal of reducing database calls per request while using the Fargate tasks' memory, with no additional managed service required.

Why this answer

Caching reference data in the application's own process memory means each request is served locally with no database or network call, which is the most direct way to reduce per-request database calls. Because the data is small, changes rarely, and tolerates slight staleness, periodic refresh is simple and effective. External caches like Redis or Memcached still require a network call, and EFS is storage rather than memory.

Exam trap

The trap here is treating any cache, such as ElastiCache, as equivalent to in-process caching, when external caches still require a network round trip per request.

265
MCQhard

Based on the exhibit, which change will most improve the CloudFront cache hit ratio for the static assets while still serving the same files to all users?

A.Create a custom cache policy that includes only the v query string and excludes cookies.
B.Enable Origin Shield and keep the current cache behavior unchanged.
C.Move the static assets to individual presigned URLs for each viewer.
D.Increase the CloudFront default TTL to 24 hours while continuing to forward all cookies and query strings.
AnswerA

This removes unnecessary cache-key fragmentation. Since all users receive identical static files, forwarding user-specific cookies and irrelevant query strings destroys cache reuse. Keeping only the version parameter preserves correct object variation while allowing many more requests to hit the same cached object at the edge.

Why this answer

The CloudFront cache hit ratio for static assets is reduced when query strings and cookies are forwarded to the origin, because each unique combination creates a separate cache entry. By creating a custom cache policy that includes only the 'v' query string (used for versioning) and excludes cookies, CloudFront can cache a single object for all users regardless of other query parameters or cookie values, maximizing cache hits while still serving the same file.

Exam trap

The trap here is that candidates assume increasing TTL or enabling Origin Shield will fix a low cache hit ratio, when the real issue is an overly broad cache key caused by forwarding all query strings and cookies.

How to eliminate wrong answers

Option B is wrong because enabling Origin Shield reduces load on the origin and improves cache fill efficiency, but it does not address the root cause of low cache hit ratio—forwarding all query strings and cookies still creates many unique cache keys. Option C is wrong because moving static assets to individual presigned URLs for each viewer would force CloudFront to treat each URL as a distinct object, drastically reducing the cache hit ratio and defeating the purpose of caching. Option D is wrong because increasing the default TTL to 24 hours while continuing to forward all cookies and query strings does not reduce the number of unique cache keys; CloudFront will still cache separate copies for each cookie and query string combination, so the cache hit ratio remains low.

266
MCQmedium

A high-volume analytics dashboard writes streaming click events that must be processed by multiple independent consumers. Which service is most appropriate?

A.Amazon Route 53
B.Amazon EBS
C.Amazon Kinesis Data Streams
D.AWS DataSync
AnswerC

Kinesis Data Streams retains an ordered, replayable record sequence for a configurable period, and each consumer reads independently via its own iterator, so multiple analytics applications process the same click events without competing for messages.

Why this answer

Amazon Kinesis Data Streams is the most appropriate service because it is designed for real-time streaming data ingestion and can be consumed by multiple independent consumers in parallel. Each shard within a Kinesis stream supports up to 5 read transactions per second and a total data read rate of 2 MB per second, allowing multiple consumer applications to process the same stream of click events concurrently without interfering with each other.

Exam trap

The trap here is that candidates often confuse Amazon Kinesis Data Streams with Amazon SQS or Amazon SNS, but SQS is a message queue for decoupled point-to-point communication and SNS is a pub/sub notification service, neither of which natively supports multiple independent consumers processing the same stream of data with replay capability.

How to eliminate wrong answers

Option A is wrong because Amazon Route 53 is a DNS web service that translates domain names to IP addresses and does not ingest or process streaming data. Option B is wrong because Amazon EBS provides block-level storage volumes for EC2 instances and cannot natively support multiple independent consumers reading a continuous stream of events. Option D is wrong because AWS DataSync is a data transfer service for moving large datasets between on-premises storage and AWS services, not for real-time streaming event processing.

267
MCQeasy

A team wants to delegate IAM management to developers, but must ensure developers can never grant themselves permissions beyond a specific limit. Which AWS mechanism best matches this requirement?

A.Use an IAM permission boundary on roles/users that developers create, so the developers’ effective permissions are capped by the boundary policy.
B.Rely only on their IAM managed policies and instruct developers to self-check against internal guidelines.
C.Use a service control policy (SCP) that applies only to the developers’ IAM users in the account.
D.Use a KMS key policy to restrict IAM actions, because IAM actions can be controlled with KMS.
AnswerA

Permission boundaries constrain the maximum permissions that an identity can receive. Even if developers attach an identity policy that allows broader actions, the effective permissions are limited to the intersection of the identity policy and the boundary.

Why this answer

IAM permission boundaries are the correct mechanism because they allow a developer to create IAM roles or users, but explicitly cap the maximum permissions those entities can have. The boundary policy acts as a ceiling, so even if a developer attaches a permissive managed policy, the effective permissions are the intersection of the boundary and the attached policy. This directly enforces the requirement that developers cannot grant themselves permissions beyond a specific limit.

Exam trap

The trap here is confusing service control policies (SCPs) with permission boundaries, as both can limit permissions, but SCPs apply account-wide and cannot be selectively applied to only developers' IAM users, while permission boundaries are attached directly to the IAM entity.

Why the other options are wrong

B

Option B relies on manual self-policing without any technical enforcement, which cannot prevent developers from granting themselves permissions beyond the specified limit. AWS IAM has no built-in mechanism to enforce internal guidelines automatically.

C

Service control policies (SCPs) apply to all IAM users and roles in an AWS account, not just to specific developers' IAM users. SCPs cannot target individual users; they apply at the account, OU, or organization level.

D

KMS key policies control access to KMS keys, not IAM actions. They cannot restrict IAM permissions or prevent developers from granting themselves elevated IAM privileges.

When would these options actually be correct?

B

This option would be correct in a scenario where the question asks for a non-technical, process-based approach to manage permissions, such as 'Which method relies on developer compliance with written policies?' or 'Which approach is suitable for a small team with high trust and no need for automated controls?'

C

If the question required restricting permissions for all principals in an AWS account (e.g., to enforce a maximum permission boundary for the entire account), an SCP would be the correct mechanism. For example: 'A company wants to ensure no IAM user or role in the account can access a specific service, regardless of IAM policies.'

D

A question that asks: 'Which mechanism restricts which IAM users can use a specific KMS key for encryption operations?' — then a KMS key policy would be correct because it directly controls access to the key.

Why candidates pick the wrong answer

B

Candidates may think that clear guidelines and self-checks are sufficient for compliance, underestimating the need for technical enforcement to prevent privilege escalation in IAM.

C

Candidates may confuse SCPs with IAM permission boundaries, thinking SCPs can be applied to individual users. They might also believe SCPs are a fine-grained control for specific IAM entities, when in reality SCPs are account-level guardrails.

D

Candidates may confuse KMS key policies with IAM policies, thinking that since KMS integrates with IAM, its key policies can also control IAM actions, or they may misremember that KMS can be used for authorization beyond encryption.

268
MCQeasy

A web application uses an Amazon Aurora DB cluster. The workload is becoming read-heavy, and the application team wants to increase read throughput without changing the database schema. They can adjust the application to route reads differently. What should they do?

A.Add Aurora read replicas and route read queries to the cluster reader endpoint
B.Switch the cluster to Multi-AZ with a longer failover target clock
C.Move all reads to the writer endpoint to reduce connection overhead
D.Disable automated backups to reduce storage overhead and speed reads
AnswerA

Aurora read replicas scale read throughput by creating up to 15 independent instances that share the same distributed storage volume. The cluster reader endpoint automatically routes new connections to any available replica, allowing SELECT queries to run in parallel across multiple instances while the writer instance focuses on update operations. Because replicas require no data copy and remain fully in sync with minimal replica lag, this approach directly addresses a read-heavy workload without schema or application refactoring.

Why this answer

Adding Aurora read replicas and routing read queries to the cluster reader endpoint is the correct approach because Aurora replicas share the same underlying storage volume as the primary instance, so they can serve read traffic with minimal replication lag. The reader endpoint automatically load-balances connections across all available replicas, increasing aggregate read throughput without requiring any schema changes.

Exam trap

The trap here is confusing Multi-AZ with read replicas: candidates often think Multi-AZ improves read performance, but in standard RDS Multi-AZ the standby is passive and cannot serve reads, whereas Aurora's architecture allows all replicas to actively handle read traffic.

How to eliminate wrong answers

Option B is wrong because Multi-AZ with a longer failover target clock does not increase read throughput; it only provides high availability by maintaining a standby in another Availability Zone, and the standby cannot serve reads. Option C is wrong because moving all reads to the writer endpoint would increase load on the single writer instance, reducing overall read throughput and potentially impacting write performance. Option D is wrong because disabling automated backups does not increase read throughput; backups are stored separately and do not affect the performance of read operations on the cluster.

269
MCQmedium

Your company currently uses an Application Load Balancer (ALB) in front of a service that receives a large number of TCP and UDP packets (including UDP-based telemetry). During load tests, you need to support both TCP and UDP traffic at high throughput while keeping stable IP endpoints for a downstream firewall allowlist. Which change best meets these requirements?

A.Switch to a Network Load Balancer (NLB) configured for TCP/UDP, and use Elastic IPs to provide stable endpoint IP addresses for allowlisting.
B.Keep the ALB and add an AWS WAF Web ACL to improve throughput and add static IP support.
C.Replace the ALB with an API Gateway REST API to support UDP because API Gateway can forward UDP packets.
D.Use an Auto Scaling group with multiple EC2 instances and no load balancer to avoid any networking bottlenecks.
AnswerA

NLB operates at Layer 4 and supports both TCP and UDP. For stable IP allowlists, you can associate Elastic IP addresses with the NLB so the load balancer exposes consistent IPs (as opposed to relying on dynamic addresses). This combination directly satisfies protocol support and stable endpoint requirements.

Why this answer

A Network Load Balancer (NLB) operates at Layer 4 and can handle both TCP and UDP traffic natively, unlike an ALB which only supports HTTP/HTTPS and cannot forward UDP packets. By assigning Elastic IPs to the NLB, you provide stable, static IP endpoints that can be added to a downstream firewall allowlist, meeting both the protocol and throughput requirements.

Exam trap

The trap here is that candidates assume an ALB can handle all traffic types because it is the most commonly used load balancer, but they forget that ALB is strictly Layer 7 and cannot process UDP packets, making the NLB the only correct choice for mixed TCP/UDP workloads requiring static IPs.

How to eliminate wrong answers

Option B is wrong because an ALB cannot handle UDP traffic (it only supports HTTP/HTTPS and WebSocket), and AWS WAF does not add static IP support or improve throughput for Layer 4 traffic. Option C is wrong because API Gateway REST APIs do not support UDP traffic; they only handle HTTP/HTTPS and WebSocket protocols. Option D is wrong because removing the load balancer eliminates the stable IP endpoint required for the firewall allowlist and introduces a single point of failure, while also not addressing the need for high-throughput TCP/UDP handling with a consistent front-end IP.

270
MCQhard

A company runs an internal analytics application on Amazon RDS for PostgreSQL. The database is used heavily from 08:00 to 18:00 on weekdays, but outside those hours it receives almost no queries. The company must keep the database available at all times and cannot tolerate downtime during business hours. The team wants to reduce the cost of running this database. Which approach is the MOST cost-effective while meeting the availability requirement?

A.Convert the DB instance to a Multi-AZ deployment and purchase a Reserved Instance for the primary instance.
B.Enable Aurora Auto Scaling with a reader endpoint and route all read queries to the reader.
C.Stop the DB instance every evening and start it again each weekday morning.
D.Migrate the workload to Aurora Serverless v2 with a minimum capacity that scales down during idle periods.
AnswerD

Aurora Serverless v2 scales capacity in fine-grained increments and can scale down to a low minimum during idle periods while remaining available to accept connections. Because the database is almost unused outside business hours but must stay online, this pay-for-what-you-use capacity model reduces cost more effectively than a fixed-size RDS instance that bills the same rate around the clock.

Why this answer

The workload has a sharp, predictable busy period and long idle windows, yet the database must remain available at all times. Aurora Serverless v2 capacity scales down during idle periods without stopping the database, so the company pays far less during the quiet hours while preserving continuous availability, which is the most cost-effective option that satisfies the constraint.

Exam trap

The trap here is treating any idle period as an opportunity to stop the database, when the requirement for continuous availability rules out stopping and favors a capacity model that scales down while staying online.

271
MCQmedium

A company runs a two-tier web application on Amazon EC2 instances in a public subnet. The EC2 instances must access an Amazon Aurora MySQL DB cluster in private subnets. A security engineer must ensure that only the web tier can reach the database on port 3306, and that no other resources in the VPC can connect. Which combination of security group configuration and subnet placement should the engineer implement?

A.Attach a security group to the Aurora cluster that allows inbound TCP 3306 from the CIDR block of the public subnet, and place the Aurora cluster in private subnets.
B.Attach a security group to the Aurora cluster that allows inbound TCP 3306 from the security group attached to the EC2 instances, and place the Aurora cluster in the same public subnet as the EC2 instances.
C.Attach a network ACL to the private subnets that allows inbound TCP 3306 from the public subnet CIDR block, and rely on the default security group for the Aurora cluster.
D.Attach a security group to the Aurora cluster that allows inbound TCP 3306 from the security group attached to the EC2 instances, and place the Aurora cluster in private subnets.
AnswerD

Referencing the web tier's security group as the source in the database security group's inbound rule allows any instance that carries that security group to connect on 3306, regardless of its private IP address. Placing Aurora in private subnets removes any route to the internet, so only resources inside the VPC can attempt a connection, satisfying the least-privilege requirement.

Why this answer

The secure pattern is to keep the database in private subnets and use a security group inbound rule that references the web tier's security group as the source. Security group references are evaluated dynamically, so only instances carrying that group can open a connection on the database port, and private subnet placement prevents any inbound path from the internet.

Exam trap

The trap here is assuming that specifying the public subnet CIDR block in the database security group is equivalent to allowing only the web tier, when in fact a CIDR rule permits every resource in that subnet.

272
MCQmedium

A mobile app reads the same product catalog items repeatedly throughout the day. The DynamoDB table is already properly keyed, but read latency is still a problem during sales events. The team can tolerate eventually consistent reads and wants the least disruptive change. What should they add?

A.Add a global secondary index for every frequently viewed product attribute.
B.Enable DynamoDB Accelerator to cache frequently accessed items in memory.
C.Switch the table to on-demand capacity mode to reduce latency.
D.Move the catalog to Aurora and use a read replica for every region.
AnswerB

DynamoDB Accelerator, or DAX, is the best fit for repeated reads of the same items when eventual consistency is acceptable. It provides an in-memory cache in front of DynamoDB and can dramatically reduce read latency for hot catalog items during traffic spikes. Because the table schema is already sound, DAX adds performance without forcing a redesign of keys or access patterns.

Why this answer

DynamoDB Accelerator (DAX) is a fully managed, in-memory cache that reduces read latency for frequently accessed items by orders of magnitude, from single-digit milliseconds to microseconds. Since the team can tolerate eventually consistent reads, DAX is ideal because it caches read results and serves them without additional DynamoDB read capacity consumption, making it the least disruptive change — no schema changes or application rewrites are required.

Exam trap

The trap here is that candidates often confuse throughput scaling (on-demand capacity) with latency reduction, or they over-engineer the solution by migrating to a different database when a simple caching layer (DAX) is the least disruptive and most cost-effective fix.

Why the other options are wrong

A

Adding a GSI for every frequently viewed attribute does not reduce read latency for repeated reads of the same items; it adds storage and write costs without addressing the latency caused by repeated reads from disk.

C

Switching to on-demand capacity mode addresses throughput provisioning, not read latency. Latency issues from repeated reads are better solved by caching, not capacity mode changes.

D

Moving to Aurora and using read replicas is a much more disruptive change than enabling DAX, and it does not address the core issue of caching frequently accessed items in memory for low-latency reads. Aurora is a relational database, not a key-value store like DynamoDB, and the question specifies the team wants the least disruptive change.

When would these options actually be correct?

A

A question where the app needs to query items by non-key attributes (e.g., filtering or sorting by product category) and the current table key does not support those access patterns efficiently.

C

A DynamoDB table experiences frequent throttling errors due to unpredictable traffic spikes, and the team wants to avoid manual capacity management. On-demand capacity mode would be correct to automatically scale throughput.

D

This option would be correct in a scenario where the application requires complex SQL queries, joins, or transactions that DynamoDB cannot support, and the team is already considering migrating to a relational database. For example: 'A company needs to run complex analytical queries on product catalog data and requires high availability across multiple regions.'

Why candidates pick the wrong answer

A

Candidates may think that indexing more attributes will speed up reads, but GSIs are for alternative query patterns, not for caching repeated reads of the same items.

C

Candidates may confuse throughput provisioning with latency optimization, assuming that 'on-demand' automatically reduces latency by scaling instantly.

D

Candidates may think that moving to a more powerful database like Aurora with read replicas will solve latency issues, especially if they are more familiar with relational databases than DynamoDB caching solutions. They might also overestimate the disruption of enabling DAX compared to a full database migration.

273
MCQmedium

A claims portal stores audit logs in S3. The compliance team requires that logs cannot be overwritten or deleted for seven years. What should be configured?

A.S3 server access logging
B.S3 versioning only
C.S3 Object Lock in compliance mode with an appropriate retention period
D.S3 lifecycle expiration after seven years
AnswerC

S3 Object Lock in compliance mode enforces a write-once-read-many (WORM) model where every object version is locked for a specified retention period. During that period, neither the object owner, the bucket owner, nor even the AWS root user can overwrite or delete the object; any such attempt fails. This makes it the only option that provides true immutability and regulatory-grade protection, such as meeting SEC Rule 17a-4 requirements, for audit log storage.

Why this answer

C is correct because S3 Object Lock in compliance mode enforces a write-once-read-many (WORM) model that prevents any user, including the root user, from overwriting or deleting objects for the specified retention period. This meets the compliance team's requirement that logs cannot be altered or removed for seven years, as compliance mode provides the highest level of protection and cannot be bypassed or shortened.

Exam trap

The trap here is that candidates often confuse versioning (which only preserves history but allows deletion via delete markers) with Object Lock's ability to enforce immutability, or they mistakenly think server access logging or lifecycle policies can prevent data modification.

How to eliminate wrong answers

Option A is wrong because S3 server access logging only records requests made to the bucket (audit trail), but does not prevent overwrites or deletions of existing objects. Option B is wrong because S3 versioning alone preserves previous versions of objects but does not prevent deletion of the current version or overwriting of object data; a delete marker can still be placed, and objects can be permanently deleted if versioning is suspended. Option D is wrong because S3 lifecycle expiration after seven years would automatically delete objects after that period, but it does not prevent premature deletion or overwriting before the seven-year mark.

274
MCQmedium

A media company runs a nightly batch job that processes video thumbnails. The batch can be interrupted at any time, and workers can resume automatically from checkpoints (a termination does not corrupt progress). The business goal is the lowest possible compute cost, and occasional interruptions are acceptable as long as the job continues automatically. Which approach is most cost-optimized?

A.Run the job on On-Demand EC2 instances to avoid interruptions
B.Use EC2 Spot Instances and implement interruption handling with checkpoint-based restarts
C.Buy Reserved Instances for the entire job window because interruptions are acceptable anyway
D.Use Savings Plans but schedule the job only during business hours to reduce the commit cost
AnswerB

Spot Instances bill at up to a 90% discount versus On-Demand, and the checkpoint-based restart design means a two-minute interruption notice causes no lost work. Since the stem accepts interruptions provided the job resumes automatically, Spot delivers the lowest compute cost without violating any stated constraint.

Why this answer

Spot Instances offer the lowest compute cost (up to 90% discount vs. On-Demand) and the checkpoint-based design ensures that interruptions are handled gracefully without data loss. The job can resume automatically from the last checkpoint, making Spot Instances ideal for fault-tolerant, interruptible batch workloads.

Exam trap

The trap here is that candidates assume Reserved Instances or Savings Plans are always cheaper for predictable workloads, but they overlook that Spot Instances can be even cheaper and are perfectly suited for fault-tolerant, interruptible batch jobs without any upfront commitment.

How to eliminate wrong answers

Option A is wrong because On-Demand instances are significantly more expensive than Spot Instances, and the business explicitly accepts occasional interruptions, so paying a premium for uninterrupted compute is not cost-optimized. Option C is wrong because Reserved Instances require a 1- or 3-year commitment and are designed for steady-state workloads, not for a nightly batch job that can be interrupted; the cost savings are less than Spot and the commitment is unnecessary. Option D is wrong because Savings Plans also require a commitment (1 or 3 years) and scheduling the job only during business hours does not reduce the commit cost; the job runs nightly, so this approach would either waste committed spend or require overprovisioning, making it less cost-effective than Spot.

275
MCQmedium

A data engineering team runs a nightly ETL job on EC2. The job can be checkpointed every 5 minutes and can be retried from the last checkpoint if the instance terminates. The job runtime varies from 2 to 4 hours, and the team has no need for a specific instance type, as long as it completes before 7:00 AM local time. They currently run the job on On-Demand EC2, leading to high monthly compute cost. Which change best reduces cost while maintaining the business deadline?

A.Use Spot Instances for the ETL workload, and configure the job to checkpoint frequently and restart on interruption.
B.Use Reserved Instances with a 1-year term to lower costs, since reservations provide discounts for any usage.
C.Switch to On-Demand but enable Auto Scaling so the job finishes faster during peak hours.
D.Use Spot Instances but disable checkpointing to simplify the application.
AnswerA

Spot Instances cost substantially less than On-Demand, and the five-minute checkpointing means an interruption loses at most five minutes of work. Restarting from the last checkpoint keeps the 2–4 hour job within the 07:00 deadline, unlike uncheckpointed interruption.

Why this answer

Spot Instances offer significant cost savings (up to 90%) compared to On-Demand, and the ETL job's ability to checkpoint every 5 minutes and restart from the last checkpoint makes it resilient to Spot interruptions. This allows the team to meet the 7:00 AM deadline while drastically reducing compute costs, as the job can be retried on new Spot capacity if interrupted.

Exam trap

The trap here is that candidates may overlook the checkpointing requirement and choose Reserved Instances (B) thinking they always reduce costs, or disable checkpointing (D) assuming simplicity is better, without realizing that Spot Instances require fault tolerance to be cost-effective.

Why the other options are wrong

B

Reserved Instances require a 1-year commitment and are cost-effective only for steady-state, predictable workloads. The nightly ETL job runs only 2-4 hours per day, so the discount does not offset the cost of paying for 24/7 reserved capacity, making it more expensive than Spot Instances.

C

Auto Scaling does not reduce cost; it adds more instances, increasing cost. The job already runs within the deadline, so scaling out is unnecessary and more expensive.

D

Disabling checkpointing removes the ability to resume from the last checkpoint on interruption, which is critical for Spot Instances that can be terminated at any time. Without checkpointing, the job would have to restart from scratch, likely missing the 7:00 AM deadline.

When would these options actually be correct?

B

A company runs a 24/7 web server fleet with consistent baseline usage. They need to reduce costs for the always-on instances and can commit to a 1-year or 3-year term. Reserved Instances would provide a significant discount over On-Demand for this steady-state workload.

C

A question where the ETL job is at risk of missing a tight deadline due to variable runtime, and cost is not the primary concern. For example: 'A batch job must complete within 1 hour, but often takes 90 minutes on a single instance. Which change ensures it finishes on time?'

D

If the ETL job were idempotent and very short (e.g., under 5 minutes), or if checkpointing introduced unacceptable overhead, then disabling it might be acceptable. For example, a simple data transformation that runs in under 5 minutes and can be safely restarted without data loss.

Why candidates pick the wrong answer

B

Candidates know Reserved Instances offer discounts and may assume any usage benefits, overlooking that the discount applies only to the reserved capacity, which is wasted when the instance is idle for most of the day.

C

Candidates may think Auto Scaling always reduces cost by optimizing resource usage, but here it would increase cost by adding instances without need.

D

Candidates may think simplifying the application by removing checkpointing reduces complexity and overhead, not realizing that checkpointing is essential for fault tolerance with Spot Instances to meet deadlines.

276
MCQmedium

A healthcare company needs to store patient records in Amazon DynamoDB. The records must be highly available and durable across multiple Availability Zones. The company also requires the ability to recover the table to any point in time within the last 35 days in case of accidental writes or deletions. Which solution meets these requirements?

A.Create a DynamoDB table and enable point-in-time recovery (PITR).
B.Create a DynamoDB table with a read replica in another AWS Region and enable point-in-time recovery (PITR).
C.Create a DynamoDB table with a global secondary index and enable point-in-time recovery (PITR).
D.Create a DynamoDB table with on-demand capacity mode and enable point-in-time recovery (PITR).
AnswerA

DynamoDB automatically replicates data across multiple Availability Zones within an AWS Region, providing high availability and durability. Enabling point-in-time recovery (PITR) allows restoration to any point in time within the last 35 days, protecting against accidental writes or deletions. This solution meets both the durability and recovery requirements without additional complexity.

Why this answer

DynamoDB tables are automatically replicated across multiple Availability Zones within a Region, ensuring high availability and durability. Enabling point-in-time recovery (PITR) provides continuous backups and allows restoration to any point in time within the last 35 days. This combination meets the requirements without additional configuration such as global tables or secondary indexes.

Exam trap

The trap here is thinking that additional features like global secondary indexes or cross-Region replicas are needed for multi-AZ durability, when DynamoDB already provides it by default.

277
MCQhard

A media company serves video thumbnails from an Amazon S3 bucket in us-east-1 to viewers across Europe and Asia. The thumbnails are immutable after upload and are requested repeatedly by the same users. The company wants to reduce latency for the global audience and reduce data transfer costs, and it does not want to modify application code. Which solution meets these requirements with the LEAST operational effort?

A.Replicate the bucket to eu-west-1 and ap-southeast-1 using S3 Cross-Region Replication, and have the application select the nearest bucket.
B.Create an Amazon CloudFront distribution with the S3 bucket as origin, enable caching, and configure an origin access control.
C.Attach an S3 bucket policy that allows public read and place an Application Load Balancer in front of the bucket.
D.Enable S3 Transfer Acceleration on the bucket and update the application to use the accelerated endpoint.
AnswerB

CloudFront caches immutable thumbnails at edge locations close to European and Asian viewers, so repeat requests are served from the edge with lower latency and fewer origin fetches, which reduces S3 data transfer cost. Origin access control restricts direct bucket access while keeping the same object URLs behind the distribution. This requires no application code change and is the least-effort global acceleration option.

Why this answer

CloudFront is the AWS content delivery network that caches objects at edge locations worldwide, so repeated thumbnail requests from Europe and Asia are served close to viewers instead of from us-east-1. Because the objects are immutable, cache hit ratios stay high, cutting origin fetches and S3 egress charges. Origin access control keeps the bucket private, and the application continues to use the distribution domain without code changes.

Exam trap

The trap here is confusing S3 Transfer Acceleration, which accelerates individual transfers to a bucket, with edge caching, which is what actually serves repeat reads quickly to a global audience.

278
MCQmedium

A company stores critical documents in an Amazon S3 bucket in the us-east-1 Region. The documents must survive an unlikely loss of the entire us-east-1 Region. The company wants a recovery point objective (RPO) of 15 minutes and a recovery time objective (RTO) of 1 hour. What should the solutions architect recommend?

A.Enable S3 server access logging and store the logs in a separate bucket in us-east-1.
B.Configure an S3 Lifecycle rule to transition objects to S3 Glacier Deep Archive after 30 days.
C.Enable S3 Transfer Acceleration on the us-east-1 bucket to speed up uploads from global clients.
D.Enable S3 Cross-Region Replication (CRR) from the us-east-1 bucket to a bucket in us-west-2 and enable S3 Versioning on both buckets.
AnswerD

S3 Cross-Region Replication asynchronously copies objects to a bucket in another Region, and versioning is required for replication. With typical replication times well under 15 minutes, this meets the RPO and provides a target bucket that can be used within the 1-hour RTO.

Why this answer

S3 Cross-Region Replication copies objects to a bucket in a different Region, and versioning is a prerequisite. Replication is asynchronous but typically completes well within a 15-minute RPO, and the destination bucket can serve as the recovery source within the 1-hour RTO.

Exam trap

The trap here is confusing performance features such as Transfer Acceleration or storage-class transitions with actual cross-Region data replication.

279
MCQmedium

A healthcare company runs a patient portal on Amazon EC2 instances behind an Application Load Balancer across two Availability Zones. A new compliance rule requires that if an entire Availability Zone fails, the portal must remain available with no manual intervention. The EC2 instances are stateless and store no session data. Which design change should the architect implement to meet this requirement?

A.Enable an Auto Scaling group that spans both Availability Zones and configure the load balancer health checks to deregister unhealthy instances.
B.Convert the EC2 instances to a single larger instance type and enable detailed monitoring in Amazon CloudWatch.
C.Configure the Application Load Balancer to use cross-zone load balancing and attach an Elastic IP address to each EC2 instance.
D.Place the EC2 instances in a single Availability Zone and create an Amazon EBS snapshot schedule every hour.
AnswerA

An Auto Scaling group spanning both Availability Zones automatically replaces instances in the surviving AZ when one AZ fails, and ALB health checks remove failed targets. This provides resilience without manual intervention, directly satisfying the compliance requirement for AZ failure tolerance.

Why this answer

High availability across Availability Zones requires compute capacity that can be automatically replaced when an AZ fails. An Auto Scaling group spanning multiple AZs, combined with load balancer health checks, ensures that unhealthy instances are removed and new instances are launched in a surviving AZ without human action.

Exam trap

The trap here is assuming that cross-zone load balancing or a larger instance type provides AZ-level resilience, when only an Auto Scaling group spanning multiple AZs can automatically replace lost capacity.

280
MCQeasy

A retail company runs a read-heavy product catalog on Amazon RDS for MySQL. During flash sales, read replicas lag behind the primary and the application serves stale prices. The team wants to scale read traffic while ensuring the application reads the most current data for price lookups. Which solution should a solutions architect recommend?

A.Enable Multi-AZ on the RDS for MySQL instance and point all reads to the standby.
B.Increase the size of the read replicas and enable automatic scaling of replicas.
C.Configure the application to retry reads on a replica until the price matches the primary.
D.Route price lookups to the primary DB instance and other reads to read replicas.
AnswerD

Read replicas are asynchronous, so they can lag during heavy write periods. Directing price lookups, which require the latest committed data, to the primary instance guarantees current values, while offloading less sensitive reads to replicas preserves read scaling. This matches the requirement for fresh prices without abandoning replica-based scaling.

Why this answer

RDS read replicas use asynchronous replication, so they can lag under write-heavy flash-sale conditions. Price lookups need the latest committed value, so sending them to the primary guarantees freshness, while other reads continue to use replicas for scale. Multi-AZ standbys cannot serve reads, and simply adding replica capacity does not eliminate lag, so routing consistency-sensitive queries to the primary is the correct design.

Exam trap

The trap here is assuming that scaling read replicas also guarantees fresh data, when asynchronous replication means replicas can lag regardless of their size.

281
MCQeasy

A retail analytics app uses Amazon RDS for PostgreSQL. Read traffic is growing, and the database CPU spikes mainly due to SELECT-heavy workloads. Writes are less frequent, and the app can tolerate eventually consistent reads for the reports. What is the most appropriate AWS-native way to improve read performance with minimal application changes?

A.Create an RDS read replica and point the reporting queries to the replica endpoint.
B.Switch the cluster to DynamoDB without redesigning the data model.
C.Enable S3 event notifications to trigger a Lambda function after each write to the database.
D.Replace the RDS instance class with a smaller size to reduce cost and improve performance.
AnswerA

Amazon RDS read replicas use asynchronous replication from the primary DB instance to one or more read-only copies, typically within the same region or cross-region. By pointing reporting and analytics queries to a replica's DNS endpoint, you offload SELECT-heavy traffic from the primary, reducing CPU and I/O contention on the source instance. This lets the primary focus on OLTP writes while analysts query near-real-time data from the replica, requiring no application schema changes. For a retail analytics app with read pressure, this is the minimal-risk, AWS-native fix.

Why this answer

Creating an RDS read replica is the most appropriate AWS-native solution because it offloads SELECT-heavy workloads from the primary database instance to a read-only copy, reducing CPU spikes on the primary. The application can tolerate eventually consistent reads for reports, which is exactly the consistency model of RDS read replicas (typically sub-second replication lag). This requires minimal application changes—only updating the reporting queries to point to the replica endpoint—and fully leverages PostgreSQL's built-in replication capabilities.

Exam trap

The trap here is that candidates might assume read replicas require application changes to handle eventual consistency, but the question explicitly states the app can tolerate eventually consistent reads, making the replica endpoint swap a minimal-change solution.

Why the other options are wrong

B

Switching to DynamoDB without redesigning the data model is not feasible because RDS PostgreSQL and DynamoDB have fundamentally different data models (relational vs. NoSQL), requiring significant application changes to adapt queries, schema, and consistency models.

C

This option does not directly improve read performance for SELECT-heavy workloads on RDS PostgreSQL. It introduces asynchronous S3 event notifications and Lambda processing, which adds latency and complexity without offloading read queries from the primary database.

D

Replacing the RDS instance with a smaller size would reduce CPU capacity, worsening performance under SELECT-heavy workloads, not improving it. The question asks for improved read performance, not cost reduction.

When would these options actually be correct?

B

A question where the application requires a fully managed NoSQL database with single-digit millisecond latency at any scale, and the team is willing to redesign the data model and application code to fit DynamoDB's key-value and document structures.

C

This option would be correct in a scenario where the requirement is to offload heavy write processing or to trigger downstream actions (e.g., data export, analytics pipeline) after each database write, and the application can tolerate eventual consistency for those downstream tasks.

D

This option would be correct in a scenario where the database is over-provisioned for the actual workload (e.g., consistently low CPU usage) and the goal is to reduce costs without negatively impacting performance. For example, a question stating 'The database CPU utilization is below 10% and costs must be minimized.'

Why candidates pick the wrong answer

B

Candidates may think DynamoDB is a universal performance solution for all read-heavy workloads, overlooking the critical need for data model compatibility and the effort required to migrate from a relational database.

C

Candidates may think that using serverless components like S3 and Lambda can scale reads, but they overlook that the bottleneck is database CPU from SELECT queries, which this option does not address.

D

Candidates may think that a smaller instance class reduces cost and assume performance is tied to cost, or they misinterpret 'improve performance' as 'reduce unnecessary resource waste'.

282
MCQhard

A patient portal must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The team wants the control to be enforceable during normal operations.

A.Use an in-memory queue on one EC2 instance
B.Use UDP messages sent directly to workers
C.Use Amazon SQS standard queue and design consumers to be idempotent
D.Use CloudFront signed URLs
AnswerC

Amazon SQS standard queues guarantee at-least-once delivery, so no event is lost, and duplicates are possible; consumers must therefore be idempotent. This satisfies the stated requirement that duplicate processing is acceptable if consumers handle idempotency.

Why this answer

Amazon SQS standard queues provide at-least-once delivery, ensuring every event is processed at least once, which matches the requirement. Duplicate processing is acceptable because the team can design consumers to be idempotent, handling duplicates without side effects. SQS is a fully managed, scalable, and durable service that enforces this behavior during normal operations without requiring custom infrastructure.

Exam trap

The trap here is that candidates may confuse 'at-least-once' delivery with 'exactly-once' delivery, or incorrectly assume that UDP or in-memory queues can provide reliable event processing, when in fact only a managed queue service like SQS with idempotent consumers meets the stated requirement for enforceability during normal operations.

How to eliminate wrong answers

Option A is wrong because an in-memory queue on a single EC2 instance is not durable, cannot survive instance failures, and does not provide at-least-once delivery guarantees across restarts or scaling events. Option B is wrong because UDP is a connectionless, unreliable protocol that does not guarantee message delivery, order, or duplicate detection, making it unsuitable for at-least-once processing. Option D is wrong because CloudFront signed URLs are used for access control to content delivery, not for event processing or messaging, and they do not provide any delivery guarantee or queue semantics.

283
MCQmedium

A team accidentally updates critical rows in an Amazon RDS for PostgreSQL database. Automated backups are enabled. They need to recover the data to the exact state as of 90 minutes ago. They also cannot risk interrupting the current production database instance while investigators validate the restored data. Which recovery strategy best meets these constraints?

A.Use point-in-time recovery (PITR) to restore to a new RDS DB instance as of 90 minutes ago, then validate and cut over after approval.
B.Restore a manual snapshot and overwrite the existing production DB instance so the data matches exactly 90 minutes ago.
C.Wait for the next automated backup window and then restart the current DB instance to roll back changes automatically.
D.Use cross-region read replicas to rewind changes and promote the replica to become the writer immediately.
AnswerA

Amazon RDS point-in-time recovery (PITR) restores a new DB instance to any fraction of a second within the backup retention window by replaying transaction logs from the last automated snapshot. Launching a separate instance preserves the production database untouched, allowing you to validate the restored data at 90 minutes ago before promoting it and updating the application connection string. After approval, you can cut over by renaming the instances or changing the DNS endpoint, minimizing downtime and risk.

Why this answer

Point-in-time recovery (PITR) for Amazon RDS allows you to restore a DB instance to any second within the backup retention period, using automated backups and transaction logs. By restoring to a new RDS instance as of 90 minutes ago, you create an isolated copy for validation without affecting the production database. This meets both the recovery point objective (RPO) of 90 minutes and the constraint of no interruption to the current production instance.

Exam trap

The trap here is that candidates confuse point-in-time recovery with snapshot restoration or assume that read replicas can be used for time-based rollbacks, but only PITR provides the exact time-targeted restore without affecting the production instance.

Why the other options are wrong

B

Restoring a manual snapshot and overwriting the production DB instance would cause downtime and data loss, as it replaces the current database entirely, violating the constraint of not interrupting production while validating.

C

Waiting for the next automated backup window does not allow recovery to a specific point 90 minutes ago; automated backups are typically taken once per day and do not support rollback to an arbitrary time.

D

Cross-region read replicas do not support rewinding changes; they replicate data asynchronously and cannot restore to a specific past point in time. Promoting a replica does not roll back the database to a previous state.

When would these options actually be correct?

B

If the requirement was to restore the entire production database to a specific past state with minimal downtime and no need for separate validation, and if overwriting the existing instance was acceptable, then restoring a manual snapshot would be appropriate.

C

If the question stated that the database needs to be restored to the state at the time of the last automated backup (e.g., 'recover to the most recent backup') and the team is willing to accept data loss from changes after that backup, then waiting for the next backup window and restarting could be a valid approach.

D

A company needs to minimize read latency for a global user base and wants to offload read traffic from the primary RDS instance. Cross-region read replicas can be promoted to become standalone databases for read-heavy workloads in different regions.

Why candidates pick the wrong answer

B

Candidates may think manual snapshots are faster or more precise than PITR, or they overlook the need to avoid production interruption during validation.

C

Candidates may mistakenly believe that automated backups enable point-in-time recovery without additional steps, or that restarting a DB instance automatically rolls back changes to a previous state.

D

Candidates may confuse read replicas with point-in-time recovery capabilities, assuming replicas can be used to revert changes, or they overestimate the rollback functionality of replicas.

284
MCQeasy

A company has a steady, predictable workload that must run continuously (24/7) in a single AWS Region. The team wants the lowest cost option available for this steady usage, but also expects they may choose different EC2 instance families in the future (without re-buying compute discounts). Which AWS purchase option best meets these goals?

A.On-Demand Instances only, because they automatically adjust to future needs
B.Compute Savings Plans, committed for a 1- to 3-year term in the Region
C.Standard Reserved Instances tied to a single instance type and Availability Zone
D.EC2 Spot Instances, because they are always cheaper than savings programs
AnswerB

Compute Savings Plans provide discounted pricing in exchange for committing to a consistent hourly spend (scoped to a Region). They apply to EC2 usage and are flexible enough that you can change EC2 instance families over time while still receiving the Savings Plans discount within the commitment scope.

Why this answer

Compute Savings Plans offer the lowest cost for steady, predictable workloads while providing instance family flexibility within a Region. Unlike Reserved Instances, they automatically apply discounts to any EC2 instance family (and even Fargate/Lambda) in the chosen Region, so the company can switch instance families in the future without losing the discount. A 1- or 3-year commitment yields significant savings (up to 66%) compared to On-Demand, making it the optimal choice for this scenario.

Exam trap

The trap here is that candidates often confuse Reserved Instances (which lock instance family and AZ) with Savings Plans (which offer regional flexibility), leading them to choose Standard Reserved Instances despite the stated requirement for future instance family changes.

Why the other options are wrong

A

On-Demand Instances are the most expensive option for steady, 24/7 workloads, as they lack the discounts of committed-use plans. The question specifically asks for the lowest cost, so On-Demand does not meet that requirement.

C

Standard Reserved Instances lock you into a specific instance type and Availability Zone, which contradicts the requirement to choose different instance families in the future without re-buying compute discounts.

D

Spot Instances can be interrupted with a 2-minute notice, making them unsuitable for a steady, continuous 24/7 workload that must run without interruption.

When would these options actually be correct?

A

On-Demand Instances would be correct for a question where the workload is unpredictable, short-term, or variable, and the priority is maximum flexibility with no upfront commitment, such as for a new application with unknown usage patterns.

C

A company has a predictable, steady workload that requires a specific instance type and is willing to commit to an Availability Zone for maximum discount, and does not need flexibility to change instance families.

D

For a fault-tolerant, stateless application that can handle interruptions (e.g., batch processing, big data, or containerized workloads) and where the lowest possible compute cost is desired, Spot Instances would be the correct choice.

Why candidates pick the wrong answer

A

Candidates may think On-Demand is the simplest and most flexible choice, and they might overlook the cost savings of committed-use plans for steady workloads, focusing only on the flexibility aspect mentioned in the option.

C

Candidates may assume Reserved Instances always offer the best savings for steady workloads, overlooking the flexibility limitations that make Compute Savings Plans more suitable here.

D

Candidates may assume Spot Instances are always the cheapest option and overlook the interruption risk, focusing only on cost without considering the workload's need for continuous availability.

285
MCQmedium

A company hosts a public-facing static website and a set of downloadable software packages. Users are distributed globally, and the packages are large, so the company wants to reduce data transfer costs and improve download latency. The content changes only when a new release is published, a few times per month. Which solution should a solutions architect recommend?

A.Store the content in Amazon S3 and distribute it through an Amazon CloudFront distribution with a cache policy that sets a long time-to-live.
B.Replicate the S3 bucket to every AWS Region and return Region-specific URLs to users.
C.Move the content to Amazon EFS and mount the file system from EC2 instances in each Region.
D.Serve the website and packages from an Amazon S3 bucket in the primary Region and enable S3 Transfer Acceleration.
AnswerA

CloudFront caches objects at edge locations close to users, which lowers download latency and reduces the volume of data transferred from the S3 origin. Because the content changes only a few times per month, a long cache time-to-live means the vast majority of requests are served from the edge at the lower CloudFront data transfer rate, cutting both latency and cost.

Why this answer

For globally distributed, cacheable static content, a CDN is the standard cost and latency optimization. CloudFront caches objects at edge locations, so most requests are served without touching the S3 origin, and the edge data transfer rate is lower than direct S3 internet transfer. A long time-to-live is appropriate because the content changes only a few times per month.

Exam trap

The trap here is confusing S3 Transfer Acceleration with content delivery; acceleration optimizes the transfer path to a bucket, while a CDN caches content near the user and reduces origin egress.

286
MCQeasy

You want to protect an Application Load Balancer (ALB) from common web exploits using AWS WAF. The application is not using CloudFront. Which AWS WAF deployment scope should you choose so the WAF rules apply to the ALB?

A.Use AWS WAF regional scope (associate the web ACL with the ALB resource)
B.Use AWS WAF CloudFront (global) scope and associate the web ACL with the ALB
C.Use AWS Shield Advanced and rely on it to inspect payloads for SQL injection and XSS
D.Use security groups only, because they can detect SQL injection patterns in HTTP requests
AnswerA

ALBs are regional resources. When you protect an ALB without CloudFront, you should use the regional WAF scope and associate the web ACL directly with the ALB, so WAF can inspect incoming requests destined for that ALB.

Why this answer

AWS WAF offers two deployment scopes: regional and CloudFront (global). Since the application is using an Application Load Balancer (ALB) without CloudFront, you must choose the regional scope. This allows you to associate the web ACL directly with the ALB resource, enabling AWS WAF to inspect HTTP/HTTPS requests for common web exploits like SQL injection and cross-site scripting (XSS) at the regional endpoint.

Exam trap

The trap here is that candidates may assume AWS WAF always requires CloudFront or that Shield Advanced provides application-layer inspection, but the exam tests the specific requirement that regional WAF is the only option for ALB without CloudFront.

Why the other options are wrong

B

AWS WAF with CloudFront (global) scope can only be associated with CloudFront distributions, not with Application Load Balancers. Since the application is not using CloudFront, this scope cannot protect the ALB.

C

AWS Shield Advanced provides DDoS protection, not application-layer web exploit detection like SQL injection or XSS. It does not inspect payloads for these threats; that is the role of AWS WAF.

D

Security groups operate at the network layer (Layer 3/4) and cannot inspect application-layer payloads for SQL injection or XSS patterns; they only filter based on IP addresses, ports, and protocols.

When would these options actually be correct?

B

If the application used CloudFront as a content delivery network in front of the ALB, then using AWS WAF with CloudFront (global) scope would be correct. The web ACL would be associated with the CloudFront distribution to protect against web exploits.

C

A question asking for the best service to protect against DDoS attacks targeting an ALB, where the application is not using CloudFront, would make Shield Advanced the correct answer.

D

If the question asked for a network-layer defense to block specific IP addresses or ports for an ALB, security groups would be the correct answer. For example: 'Which AWS service can be used to restrict inbound traffic to an ALB based on source IP address?'

Why candidates pick the wrong answer

B

Candidates may mistakenly think that AWS WAF's global scope can be applied to any AWS resource, or they may confuse the regional and global scopes, assuming the global scope is more powerful and can protect ALBs directly.

C

Candidates may confuse Shield Advanced with WAF, assuming it includes web exploit detection, or overestimate its capabilities due to its 'Advanced' naming.

D

Candidates may mistakenly believe security groups provide application-layer inspection because they are familiar with them as a primary security mechanism, overlooking that WAF is required for Layer 7 threats.

287
Multi-Selecthard

A solutions architect is optimizing the cost of a serverless data-processing pipeline. The pipeline uses AWS Lambda functions that process messages from an Amazon SQS queue and write results to Amazon DynamoDB. The team observes that Lambda invocations spike unpredictably, DynamoDB is provisioned with high capacity that is often idle, and the SQS queue occasionally accumulates a large backlog. Which two changes will most directly reduce cost while preserving the pipeline's ability to handle bursts? (Choose two.)

Select 2 answers
A.Increase the Lambda function's memory allocation to shorten execution time.
B.Configure the Lambda function's reserved concurrency to a low fixed value to cap scaling.
C.Move the SQS queue to a FIFO queue to improve ordering and throughput.
D.Switch the DynamoDB table from provisioned capacity to on-demand capacity mode.
E.Enable SQS long polling and batch multiple messages per Lambda invocation.
AnswersD, E

On-demand capacity mode charges only for the read and write requests the table actually serves, so idle provisioned capacity no longer incurs cost. Because the pipeline's traffic is unpredictable and bursty, on-demand absorbs spikes without pre-provisioning, directly aligning spend with usage and eliminating charges for capacity that sits unused during quiet periods.

Why this answer

The two most direct cost levers are matching DynamoDB billing to actual traffic by using on-demand capacity, and reducing Lambda invocation count by batching SQS messages with long polling. Together they eliminate idle provisioned capacity and cut per-invocation charges, while both changes preserve the ability to absorb unpredictable bursts without throttling the pipeline.

Exam trap

The trap here is treating Lambda tuning such as memory or reserved concurrency as a cost fix, when the predictable savings come from aligning DynamoDB billing to usage and cutting the number of billed invocations.

288
MCQhard

A mobile banking backend uses Amazon RDS for PostgreSQL. Application credentials must not be stored on the EC2 instances, and authentication should use short-lived credentials. What should the architect recommend? The design must avoid adding custom operational scripts.

A.Store the database password in user data
B.IAM database authentication for RDS with an EC2 instance role
C.Use a security group rule that allows only application instances
D.Embed the database password in the AMI
AnswerB

IAM database authentication for RDS with an EC2 instance role is the correct approach because it lets the application generate a short-lived (15-minute) authentication token using SigV4, eliminating any stored database password. The EC2 instance assumes an instance role with rds-db:connect permissions, and the application presents the token to PostgreSQL over SSL; RDS validates the token against IAM rather than a static password. This provides centralized credential management via IAM roles, automatic rotation of credentials, and ensures no secret is baked into configurations, user data, or code.

Why this answer

IAM database authentication for RDS allows EC2 instances to authenticate to PostgreSQL using a short-lived token generated via the IAM instance profile, eliminating the need to store credentials on the instance. The token is obtained by calling the RDS generate_db_auth_token API with the instance's IAM role, and it is valid for 15 minutes by default. This approach satisfies the requirement for short-lived credentials and avoids custom operational scripts.

Exam trap

The trap here is that candidates often confuse network-level controls (security groups) with authentication mechanisms, or assume that storing credentials in user data or AMIs is acceptable because they are 'hidden' from the OS, when in fact they are still long-lived and accessible via metadata or AMI inspection.

How to eliminate wrong answers

Option A is wrong because storing the database password in user data leaves it in plaintext on the instance metadata, which is accessible to any process or user with access to the instance, and it does not provide short-lived credentials. Option C is wrong because security group rules only control network access at the transport layer; they do not handle authentication or credential management, so credentials would still need to be stored on the instance. Option D is wrong because embedding the database password in the AMI hard-codes a long-lived credential into the image, which violates the requirement to avoid storing credentials on EC2 and does not provide short-lived credentials.

289
MCQmedium

A team runs an application on Amazon EC2 that connects to an Aurora database. The database password must rotate automatically every 30 days, and the application should retrieve the current secret at runtime using an IAM role. Which AWS service is the best fit?

A.AWS Systems Manager Parameter Store standard parameters.
B.AWS Secrets Manager with rotation enabled.
C.AWS KMS, because KMS stores credentials and rotates them automatically.
D.Amazon S3 with server-side encryption and versioning.
AnswerB

Secrets Manager is designed for secure secret storage with built-in rotation support and fine-grained access through IAM. In this case, the application can retrieve the current database credentials at runtime with its EC2 role, while the secret is rotated on a schedule without embedding passwords in code. This reduces operational risk, improves auditability, and avoids manual password changes that often cause outages.

Why this answer

AWS Secrets Manager is the best fit because it natively supports automatic rotation of database credentials on a schedule (e.g., every 30 days) and integrates directly with Amazon RDS/Aurora to update the password. The application can retrieve the current secret at runtime using an IAM role attached to the EC2 instance, without hardcoding credentials. Secrets Manager also provides built-in secret rotation with Lambda, ensuring zero downtime during password changes.

Exam trap

The trap here is that candidates often confuse AWS Systems Manager Parameter Store (which can store secrets but lacks native rotation) with Secrets Manager, or incorrectly assume KMS can store and rotate credentials because it handles encryption keys.

Why the other options are wrong

A

Systems Manager Parameter Store standard parameters do not support automatic rotation of secrets; they require manual updates or custom automation, whereas the question mandates automatic rotation every 30 days.

C

AWS KMS is a key management service for encryption keys, not a service for storing or rotating database passwords. It does not provide automatic rotation of secrets or direct retrieval by applications via IAM roles.

D

Amazon S3 with server-side encryption and versioning does not provide automatic password rotation or native integration with IAM roles for runtime secret retrieval; it is designed for object storage, not dynamic secrets management.

When would these options actually be correct?

A

If the question required storing configuration data (e.g., database endpoint, port) without rotation, or if the application needed to retrieve parameters at runtime without automatic rotation, Parameter Store would be appropriate.

C

A question asks which service should be used to encrypt data at rest for an application that stores sensitive files in Amazon S3, with requirements for automatic key rotation and centralized key management. AWS KMS would be the correct answer.

D

An application needs to store and retrieve large configuration files (e.g., database connection strings) with versioning and encryption, but does not require automatic rotation or IAM-based access; S3 with SSE and versioning would be appropriate.

Why candidates pick the wrong answer

A

Candidates may confuse Parameter Store with Secrets Manager because both can store secrets securely, but they overlook that Parameter Store lacks built-in rotation capabilities.

C

Candidates may confuse KMS's key rotation capability with secret rotation, or mistakenly think KMS can store credentials because it manages encryption keys that protect secrets.

D

Candidates may think S3 can store any data securely and versioning provides a form of rotation, overlooking that Secrets Manager is purpose-built for managing secrets with rotation and IAM integration.

290
Matchinghard

Match each database availability event to the AWS failover behavior that best describes it.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

The standby in another Availability Zone is promoted, and the same database endpoint remains in use after a brief reconnect.

Aurora promotes another healthy instance to writer while the shared storage layer stays intact across Availability Zones.

A manual failover can be triggered so the standby becomes primary before the reboot finishes.

Only that reader is removed from the reader set; the cluster can still serve read traffic through the remaining healthy readers.

Why these pairings

Multi-AZ RDS automatically fails over to standby; read replicas require manual redirect; Aurora uses replicas for failover; without replicas, Aurora recovers in-place.

291
MCQhard

Based on the exhibit, the company wants to lower CloudWatch and EC2 monitoring costs. Auditors require logs to be retained for 90 days, but operations only uses detailed per-instance metrics during rare troubleshooting events. Which change best reduces recurring cost while preserving the required visibility?

A.Disable CloudWatch Logs entirely and rely on application local files for 90 days.
B.Increase the number of CloudWatch alarms so that metrics are collected less expensively.
C.Set CloudWatch Logs retention to 90 days for all log groups, and switch EC2 monitoring from detailed to basic except during incidents.
D.Export all logs to Amazon S3 immediately and keep detailed monitoring enabled on every instance.
AnswerC

This directly addresses the two visible recurring cost drivers. Applying a 90-day retention policy stops indefinite log storage growth while still meeting the audit requirement. Basic monitoring is sufficient when 1-minute metrics are not required all the time, and detailed monitoring can be enabled selectively during incidents instead of paying for it across all 200 instances continuously.

Why this answer

It directly addresses the two cost drivers: CloudWatch Logs storage costs are minimized by setting a 90-day retention policy (matching the audit requirement), and EC2 detailed monitoring (1-minute metrics) is replaced with basic monitoring (5-minute metrics) during normal operations, with the ability to switch back to detailed only when needed for troubleshooting. This preserves the required log retention and the ability to obtain high-resolution metrics on demand, while eliminating the recurring cost of storing logs indefinitely and paying for detailed monitoring on every instance.

Exam trap

The trap here is that candidates may think increasing alarms or exporting logs to S3 reduces costs, but they fail to recognize that detailed monitoring is a per-instance hourly charge independent of alarms, and that S3 storage and API costs can exceed CloudWatch Logs costs if not managed carefully.

How to eliminate wrong answers

Option A is wrong because disabling CloudWatch Logs entirely and relying on application local files violates the auditor's requirement for centralized, durable log retention and makes logs inaccessible if the instance fails or is terminated. Option B is wrong because increasing the number of CloudWatch alarms does not reduce metric collection costs; alarms are billed separately and do not change the underlying cost of detailed monitoring (per-instance per-minute charges). Option D is wrong because exporting logs to S3 immediately does not reduce costs—it adds S3 storage and PUT request costs—and keeping detailed monitoring enabled on every instance continues to incur the higher per-instance monitoring fee.

292
MCQeasy

A startup runs a public web application on Amazon EC2 instances behind an Application Load Balancer. The instances are in a public subnet and currently allow SSH from 0.0.0.0/0 so that engineers can troubleshoot. Auditors flagged this exposure. Engineers still need occasional shell access to the instances, and the company wants the access to be auditable per engineer without managing bastion hosts or distributing key pairs. Which solution best meets these requirements?

A.Keep the instances in the public subnet, restrict SSH to the corporate office CIDR, and distribute a shared PEM key pair to all engineers through AWS Secrets Manager.
B.Move the instances to private subnets, remove the inbound SSH rule, and grant engineers access through AWS Systems Manager Session Manager with IAM policies and session logging to Amazon S3 and CloudWatch Logs.
C.Attach an EC2 instance profile granting AmazonSSMManagedInstanceCore to the instances and open port 22 only to the VPC CIDR so Session Manager can reach the instances.
D.Deploy a bastion host in a public subnet with a security group that allows SSH only from the corporate CIDR, and have engineers forward through it to reach the instances.
AnswerB

Session Manager connects to instances through the Systems Manager agent without inbound ports or a bastion host, so the SSH rule can be deleted entirely. Access is governed by IAM, so each engineer's session is attributable, and session logging to Amazon S3 and CloudWatch Logs produces the audit trail the auditors requested without distributing or rotating SSH key pairs.

Why this answer

Session Manager removes the need for inbound SSH, bastion hosts, and key pairs by having the Systems Manager agent establish outbound connections to the service. IAM policies determine which engineers can start sessions on which instances, giving per-person attribution, and session logging to Amazon S3 and CloudWatch Logs satisfies the audit requirement without exposing any listening port.

Exam trap

The trap here is assuming Session Manager requires an open SSH port or a bastion host to function, when the agent only needs outbound connectivity to Systems Manager endpoints.

293
MCQeasy

A service performs many repeated read requests for the same DynamoDB items. The reads are latency-sensitive, but the application can tolerate slightly stale data. Which AWS service is the best fit to reduce read latency?

A.Amazon DAX (DynamoDB Accelerator)
B.Amazon S3 Select
C.Amazon SQS FIFO queue
D.AWS Lambda provisioned concurrency
AnswerA

Amazon DAX is an in-memory cache for DynamoDB. It reduces latency for repeated reads by caching results and serving subsequent read requests from the DAX cluster rather than repeatedly calling DynamoDB. Because it provides cached reads that may be slightly stale, it matches the scenario’s tolerance.

Why this answer

Amazon DAX (DynamoDB Accelerator) is an in-memory cache specifically designed for DynamoDB. It reduces read latency from single-digit milliseconds to microseconds by caching frequently accessed items, and it supports eventually consistent reads, which aligns with the application's tolerance for slightly stale data. DAX handles repeated read requests without additional DynamoDB read capacity unit consumption, making it the optimal choice for this latency-sensitive workload.

Exam trap

The trap here is that candidates often confuse caching services (DAX) with data retrieval services (S3 Select) or assume that a queue (SQS) or compute optimization (Lambda provisioned concurrency) can solve read latency issues, when only a purpose-built in-memory cache like DAX directly addresses repeated DynamoDB reads with stale data tolerance.

How to eliminate wrong answers

Option B (Amazon S3 Select) is wrong because it retrieves subsets of data from objects stored in S3 using SQL-like queries, not from DynamoDB items, and it does not provide a caching layer to reduce read latency for repeated DynamoDB reads. Option C (Amazon SQS FIFO queue) is wrong because it is a message queuing service for decoupling and ordering messages, not a caching or read-acceleration service for DynamoDB; it adds latency rather than reducing it for repeated reads. Option D (AWS Lambda provisioned concurrency) is wrong because it pre-warms Lambda execution environments to reduce cold starts, but it does not cache DynamoDB items or reduce read latency for repeated database queries.

294
Drag & Dropmedium

Order the steps to create a static website using Amazon S3 and CloudFront.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

S3 bucket with hosting, upload files, CloudFront distribution, configure CloudFront, then DNS.

295
MCQmedium

A Lambda function for a claims portal needs to read a database password. The password must rotate automatically every 30 days and should not be stored in environment variables. Which service should be used?

A.AWS Systems Manager Parameter Store SecureString without automation
B.An encrypted object in Amazon S3
C.A KMS-encrypted Lambda environment variable
D.AWS Secrets Manager with rotation enabled
AnswerD

AWS Secrets Manager with rotation enabled stores the database credentials as a secret encrypted with a KMS key and automatically invokes a rotation Lambda function on a schedule to change the password. The Lambda can call GetSecretValue to fetch the current credentials, and the managed rotation eliminates manual credential updates, making it the right choice.

Why this answer

AWS Secrets Manager is the correct choice because it is purpose-built for securely storing, automatically rotating, and managing secrets like database passwords. It supports automatic rotation every 30 days via a built-in Lambda rotation function, and it avoids storing the password in environment variables, which are visible in the Lambda console and logs.

Exam trap

The trap here is that candidates often confuse AWS Systems Manager Parameter Store SecureString with Secrets Manager, assuming Parameter Store can also handle automatic rotation, but Parameter Store lacks native rotation capabilities and requires custom automation.

How to eliminate wrong answers

Option A is wrong because AWS Systems Manager Parameter Store SecureString can store encrypted secrets but does not support automatic rotation without additional custom automation (e.g., a scheduled Lambda function). Option B is wrong because an encrypted object in Amazon S3 requires manual management of encryption keys and rotation, and the Lambda function would need to download and decrypt the object each time, adding complexity and latency. Option C is wrong because a KMS-encrypted Lambda environment variable, while encrypted at rest, is still stored as an environment variable that can be exposed in the Lambda function's configuration, logs, or error messages, and it does not support automatic rotation.

296
Multi-Selectmedium

A logistics company runs a stateless order-tracking API on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. The architect must ensure the API survives the loss of an entire Availability Zone and that unhealthy instances are replaced automatically. (Choose two.)

Select 2 answers
A.Enable termination protection on all EC2 instances so the Auto Scaling group cannot remove them during a zone failure
B.Create a second Auto Scaling group in a different Region and use Amazon Route 53 latency-based routing to distribute traffic
C.Attach an Application Load Balancer target group health check and enable ELB health checks on the Auto Scaling group so unhealthy instances are terminated and replaced
D.Place the instances in a single Availability Zone and enable detailed CloudWatch monitoring with a 1-minute granularity
E.Configure the Auto Scaling group to span at least two Availability Zones and set the desired capacity to a number that keeps instances running in each zone
AnswersC, E

Enabling Elastic Load Balancing health checks on the Auto Scaling group means the group uses the load balancer's health status rather than only EC2 status checks, so instances that fail application-level checks are replaced. This ensures automatic recovery from unhealthy instances, which is the second requirement. The target group health check defines what the load balancer considers healthy.

Why this answer

Zone-level resilience with EC2 Auto Scaling requires the group to span multiple Availability Zones with enough desired capacity to run in each, so a single zone failure leaves serving capacity. Automatic replacement of unhealthy instances requires ELB health checks on the group, which makes the group act on the load balancer's health determination rather than only EC2 status checks. Together these two settings deliver the required survival and self-healing behaviour.

Exam trap

The trap here is assuming that enabling monitoring or termination protection contributes to resilience, when only multi-AZ capacity and ELB health checks do.

297
MCQhard

A company runs a critical API on Amazon EC2 behind an Application Load Balancer in a single AWS Region. The business requires the API to keep serving traffic if an entire Availability Zone becomes unavailable, and the recovery must not depend on any manual step. The database is Amazon RDS for PostgreSQL configured as a Single-AZ instance. Which combination of changes should a solutions architect implement to meet these requirements?

A.Configure the Auto Scaling group to span at least two Availability Zones with a health check grace period, and convert the RDS instance to a Multi-AZ DB instance deployment.
B.Enable Multi-AZ on the Application Load Balancer by adding a second listener in a different Availability Zone.
C.Take hourly RDS snapshots and configure an Auto Scaling lifecycle hook to restore the snapshot into a new Availability Zone during a failure.
D.Create an RDS read replica in a second Availability Zone and update the application to write to the replica when the primary fails.
AnswerA

Spreading the Auto Scaling group across multiple Availability Zones lets the load balancer route to healthy instances when one zone fails, and Multi-AZ RDS maintains a synchronous standby in another zone with automatic failover to the same endpoint. Together these provide automatic, hands-off recovery for both compute and database tiers.

Why this answer

Automatic recovery across an Availability Zone failure needs redundant compute and a database that fails over on its own. A multi-AZ Auto Scaling group keeps the API serving through the load balancer, and Multi-AZ RDS maintains a synchronous standby that is promoted automatically to the same endpoint, so no human action is required.

Exam trap

The trap here is reaching for an RDS read replica for high availability, when replicas are asynchronous and read-only and require manual promotion.

298
MCQhard

A company runs a containerized API on Amazon ECS with AWS Fargate. Traffic is highly variable: it peaks during business hours and drops to near zero overnight. The team wants to pay only for what they use while keeping the API responsive during peaks. Which approach BEST optimizes cost for this workload?

A.Run the tasks on Fargate Spot capacity and configure the service to scale out on demand.
B.Configure Application Auto Scaling on the ECS service using target tracking on a metric such as average CPU or requests per task, with a minimum task count of one.
C.Migrate the workload to a single large EC2 instance running the containers with a 3-year Reserved Instance.
D.Provision a fixed task count sized for the highest observed peak and leave it running continuously.
AnswerB

Application Auto Scaling with target tracking adjusts the desired task count to match demand, so the service scales out during business-hour peaks and scales in overnight when traffic nears zero. Setting a low minimum keeps a task ready to absorb the first requests, preserving responsiveness. Because Fargate bills per task-second, scaling in directly reduces spend, making this the best cost-and-performance balance.

Why this answer

Variable traffic that peaks in the day and falls to near zero at night is the classic case for elastic horizontal scaling. Application Auto Scaling with target tracking grows and shrinks the ECS task count to follow demand, and because Fargate charges per task-second, scaling in overnight removes the cost of idle capacity. A low minimum keeps the API warm and responsive for the first requests of the day.

Exam trap

The trap here is treating a cheaper capacity type like Fargate Spot as the primary lever, when the workload's need for uninterrupted responsiveness makes demand-based scaling the real cost optimizer.

299
MCQhard

A SaaS provider runs a multi-tenant application on Amazon EC2 instances behind an Application Load Balancer. Tenants are identified by a subdomain, and each tenant's data is stored in a separate Amazon S3 bucket. The provider wants HTTPS with a single certificate, automatic renewal, and the ability to add new tenant subdomains without redeploying or replacing the certificate. Which solution meets these requirements?

A.Request a public certificate in AWS Certificate Manager for the apex domain and a wildcard for its subdomains, validate it with DNS, and attach it to the Application Load Balancer HTTPS listener.
B.Import a self-signed certificate covering all current tenant subdomains into AWS Certificate Manager and attach it to the Application Load Balancer listener.
C.Terminate TLS on the EC2 instances using certificates issued by AWS Private Certificate Authority, and pass traffic from the Application Load Balancer to the instances over HTTP.
D.Store the private key and certificate in AWS Secrets Manager, and configure the Application Load Balancer to retrieve and rotate the certificate at each renewal.
AnswerA

A public ACM certificate that includes the apex domain and a wildcard covers all current and future tenant subdomains, and DNS validation allows ACM to renew the certificate automatically as long as the validation records remain in place. Attaching it to the Application Load Balancer HTTPS listener provides TLS termination without redeploying the application when tenants are added.

Why this answer

A public ACM certificate that includes the apex domain and a wildcard for its subdomains covers every current and future tenant hostname. DNS validation lets ACM renew the certificate automatically while the validation records persist, and attaching the certificate to the Application Load Balancer listener provides HTTPS termination without touching the application when new tenants are onboarded.

Exam trap

The trap here is importing a certificate that covers only today's subdomains, when a wildcard certificate requested through ACM with DNS validation covers future tenants and renews automatically.

300
MCQmedium

An application encrypts data directly with AWS KMS using an encryption context. Your KMS key policy includes a condition that allows kms:Decrypt only when the encryption context contains: "purpose" = "myapp-secrets" After a deployment, decryption fails. CloudTrail shows kms:Decrypt was called, but it was denied by the key policy due to the encryption context condition. What is the best fix?

A.Update the application code to supply the correct encryption context "purpose" = "myapp-secrets" when calling decrypt (and encrypt if rotating).
B.Add kms:Decrypt to the IAM role attached to the application without changing the key policy.
C.Disable the encryption context condition in the KMS key policy to avoid future failures.
D.Rotate the KMS key immediately and re-encrypt all secrets with a different key ID.
AnswerA

The correct fix is to make the decryption call supply the exact encryption context used at encryption time, i.e., `"purpose" = "myapp-secrets"`. AWS KMS treats the encryption context as authenticated additional data (AAD): it is not stored encrypted, but it must be provided during decryption or the operation fails. If the KMS key policy condition requires `kms:EncryptionContext:purpose` to equal that value, then every decrypt request must include that context key and value to satisfy the policy. Updating the application code to consistently pass this context—both when encrypting new secrets and when decrypting existing ones—resolves the failure without weakening the key policy or forcing key rotation, and it preserves the integrity check that the context provides.

Why this answer

The decryption failure is directly caused by the application not supplying the required encryption context in the decrypt call. The KMS key policy condition explicitly requires the encryption context to include 'purpose'='myapp-secrets' for kms:Decrypt. Without this context, the request is denied regardless of IAM permissions.

Updating the application code to pass the correct encryption context during both encrypt and decrypt operations resolves the issue.

Exam trap

The trap here is that candidates may think IAM permissions alone can override key policy conditions, but KMS requires both IAM and key policy to allow an action, and conditions in the key policy are evaluated strictly.

How to eliminate wrong answers

Option B is wrong because adding kms:Decrypt to the IAM role does not override the key policy condition; KMS requires both IAM permissions and key policy to allow the action, and the key policy condition explicitly denies decryption without the correct encryption context. Option C is wrong because disabling the encryption context condition weakens security by removing a critical access control that ensures only authorized applications with the correct context can decrypt data. Option D is wrong because rotating the KMS key does not address the root cause—the encryption context mismatch—and re-encrypting with a different key ID would still fail if the application does not supply the required context.

Page 3

Page 4 of 13

Page 5