Courseiva

SAA-C03 (SAA-C03) — Questions 751–825

935 questions total · 13pages · All types, answers revealed

Page 10

Page 11 of 13

Page 12
751
MCQmedium

A batch analytics job has unpredictable DynamoDB traffic with long idle periods and occasional spikes. Which capacity mode should minimize operational overhead and avoid paying for idle provisioned capacity? The architecture review board prefers a managed AWS-native control.

A.DynamoDB on-demand capacity mode
B.Reserved capacity for maximum daily traffic
C.Provisioned capacity set for peak traffic
D.Global tables in every Region
AnswerA

DynamoDB on-demand capacity mode automatically scales read/write capacity to match actual traffic, so you never have to estimate peaks or pre-commit throughput. You pay only for the requests you actually make, which is ideal for unpredictable, spiky workloads because sudden bursts are absorbed without throttling or manual intervention. This mode eliminates both the risk of under-provisioning and the over-provisioning cost waste of fixed capacity plans, making it the most cost-effective and operationally simple choice for this scenario.

Why this answer

DynamoDB on-demand capacity mode automatically scales to handle unpredictable traffic spikes and idle periods, charging only for the reads/writes you perform. This eliminates the need to provision capacity, reducing operational overhead and avoiding costs for idle provisioned capacity, aligning with the architecture review board's preference for a managed AWS-native control.

Exam trap

The trap here is that candidates may confuse 'reserved capacity' or 'provisioned capacity' as cost-effective for spikes, but they fail to recognize that on-demand is the only mode that eliminates idle cost and operational overhead for unpredictable workloads.

How to eliminate wrong answers

Option B is wrong because reserved capacity requires upfront commitment to a specific traffic level, which doesn't suit unpredictable spikes and idle periods, and would still incur costs for unused capacity. Option C is wrong because provisioned capacity set for peak traffic would over-provision during idle periods, leading to paying for unused capacity and increased operational overhead to manage scaling. Option D is wrong because global tables are a replication feature for multi-Region data access, not a capacity mode, and they do not address cost optimization for unpredictable traffic or idle periods.

752
MCQeasy

A web application behind an Application Load Balancer (ALB) currently allows client connections over HTTP (port 80). The security policy requires all client traffic to use HTTPS. What is the best ALB change to enforce this requirement?

A.Add an HTTP listener on port 80 with a redirect action to HTTPS on port 443, and configure an HTTPS listener using an ACM certificate
B.Enable TLS only on the target group so that traffic between the ALB and targets is encrypted, even if clients connect via HTTP
C.Turn on S3 server-side encryption to ensure data is encrypted in transit from clients to the ALB
D.Remove port 80 access by removing the port 80 listener and leave only a default target group
AnswerA

Redirecting all HTTP requests to HTTPS forces clients to use TLS when they access the application. Configuring an HTTPS listener with an ACM certificate ensures the ALB terminates TLS on port 443 using a valid certificate, directly enforcing encryption in transit for client-to-ALB traffic.

Why this answer

It uses an ALB HTTP-to-HTTPS redirect action, which is the most efficient and AWS-native way to enforce HTTPS-only traffic. The HTTP listener on port 80 automatically redirects all client requests to the HTTPS listener on port 443, which terminates TLS using an ACM certificate. This approach requires no changes to client applications and ensures compliance with the security policy at the load balancer level.

Exam trap

The trap here is that candidates may think removing the HTTP listener or enabling TLS on the target group is sufficient, but the correct approach is to use a redirect action on the HTTP listener to enforce HTTPS without breaking client connectivity.

Why the other options are wrong

B

Enabling TLS on the target group only encrypts traffic between the ALB and targets, but does not enforce HTTPS between clients and the ALB, leaving client traffic unencrypted over HTTP.

C

S3 server-side encryption encrypts data at rest in Amazon S3, not data in transit. It does not affect traffic between clients and the ALB, so it cannot enforce HTTPS for client connections.

D

Removing the port 80 listener would drop HTTP requests entirely, but the requirement is to redirect HTTP to HTTPS, not to block HTTP. This would break clients that still attempt HTTP connections.

When would these options actually be correct?

B

In a scenario where the requirement is to encrypt traffic between the ALB and backend targets (e.g., to meet compliance for internal traffic), and client-to-ALB encryption is handled separately or not required.

C

This option would be correct in a scenario where the question asks for encrypting data stored in an S3 bucket, such as 'How to ensure objects in an S3 bucket are encrypted at rest?'

D

If the security policy required that all client traffic must be HTTPS and that any HTTP requests should be rejected (not redirected), then removing the HTTP listener would be correct. For example, a strict policy that no HTTP traffic is allowed at all.

Why candidates pick the wrong answer

B

Candidates may confuse the need for end-to-end encryption with the requirement to encrypt client traffic, mistakenly thinking that securing the target group alone satisfies the HTTPS requirement.

C

Candidates may confuse encryption at rest with encryption in transit, or mistakenly think S3 can be used to secure ALB traffic due to its encryption capabilities.

D

Candidates may think that simply removing HTTP access enforces HTTPS, but they overlook the need to redirect existing HTTP clients to HTTPS to maintain accessibility and user experience.

753
MCQmedium

A partner company needs read-only access to reports in an S3 bucket for a image sharing application. The partner has its own AWS account. What is the most secure scalable access pattern?

A.Make the objects public and rely on difficult-to-guess object names
B.Create an IAM user in the company account and share the access keys
C.Copy the objects to a public website bucket
D.Create a bucket policy that grants the partner role least-privilege access to the required prefix
AnswerD

A bucket policy is an S3 resource-based policy that can grant cross-account access to a specific IAM role in the partner account. By scoping the Principal to the partner's role ARN, the Action to s3:GetObject (and s3:ListBucket if needed), and the Resource to the exact prefix path, you implement least-privilege access. The partner's role must also have a corresponding identity-based policy allowing the S3 actions, but the bucket policy is what authorizes access across accounts without storing shared keys.

Why this answer

It uses a resource-based bucket policy that grants the partner's AWS account (via its IAM role) least-privilege read-only access to a specific prefix. This avoids sharing long-term credentials, leverages AWS's cross-account trust mechanism, and ensures the partner's access is controlled through their own IAM roles, which is the most secure and scalable pattern for cross-account S3 access.

Exam trap

The trap here is that candidates often choose sharing IAM access keys (Option B) because it seems simpler, but the SAA-C03 exam emphasizes using IAM roles and resource-based policies for cross-account access to avoid long-term credential management and improve security.

How to eliminate wrong answers

Option A is wrong because making objects public with 'security through obscurity' (difficult-to-guess names) is not secure; anyone who discovers the URL can access the objects, and S3 does not enforce access control based on object name guessability. Option B is wrong because creating an IAM user in the company account and sharing access keys introduces a long-term credential that must be rotated, managed, and could be leaked; it violates the principle of least privilege and does not scale across multiple partner accounts. Option C is wrong because copying objects to a public website bucket (e.g., S3 static website hosting) makes them publicly accessible over HTTP/HTTPS with no authentication, which is insecure and does not provide read-only access control for a specific partner.

754
MCQmedium

A test environment stores logs in S3. Logs are queried for 30 days, rarely accessed for one year, and then retained for compliance. What should reduce storage cost? The design must avoid adding custom operational scripts.

A.Keep all logs in S3 Standard indefinitely
B.Move all logs immediately to S3 Glacier Deep Archive
C.S3 lifecycle policy that transitions objects to lower-cost storage classes over time
D.Use EBS snapshots for the logs
AnswerC

An S3 lifecycle policy automatically transitions objects between storage classes on defined schedules and expires them, matching the 30-day query, one-year infrequent, then compliance-retention pattern. It is configuration-driven, so no custom operational scripts are added, satisfying that constraint.

Why this answer

S3 lifecycle policies automate the transition of objects between storage classes based on age, allowing logs to move from S3 Standard (for frequent querying) to S3 Standard-IA or S3 One Zone-IA (for rare access), and eventually to S3 Glacier Deep Archive (for long-term compliance retention). This reduces storage cost without custom scripts, aligning with the requirement to avoid operational overhead.

Exam trap

The trap here is that candidates may choose Option B (immediate move to Glacier Deep Archive) thinking it minimizes cost, but they overlook the 30-day query requirement, which makes S3 Standard necessary for fast retrieval, and fail to recognize that lifecycle policies provide a graduated, automated approach.

How to eliminate wrong answers

Option A is wrong because keeping all logs in S3 Standard indefinitely incurs the highest storage cost, ignoring the cost savings from transitioning to lower-cost classes for rarely accessed and compliance-retained data. Option B is wrong because moving all logs immediately to S3 Glacier Deep Archive prevents the 30-day querying requirement, as retrieval times are hours and costs are high for frequent access, violating the design need for queryability. Option D is wrong because EBS snapshots are block-level backups for EC2 instances, not designed for log storage in S3, and would introduce unnecessary complexity and cost without addressing the tiered access pattern.

755
Multi-Selecthard

A log archive has old unattached EBS volumes and many stale snapshots. Which two actions reduce storage cost without affecting running instances? The design must avoid adding custom operational scripts.

Select 2 answers
A.Stop all EC2 instances in the account
B.Disable CloudTrail logging
C.Delete unattached EBS volumes after verifying they are no longer needed
D.Apply snapshot lifecycle policies to expire obsolete snapshots
AnswersC, D

Unattached EBS volumes in the 'available' state still incur full storage charges because EBS billing is based on provisioned capacity, not on whether the volume is attached to an instance. Before deleting, you should verify the volume is no longer needed (e.g., it is not the boot volume of a stopped instance or referenced in a launch template). You can also create a final snapshot for backup, but deleting the volume itself is what stops recurring costs.

Why this answer

Deleting unattached EBS volumes eliminates storage costs for volumes that are not in use, and since they are not attached to any running instance, this action does not affect running instances. Option D is correct because applying snapshot lifecycle policies automates the deletion of obsolete snapshots, reducing storage costs without requiring custom scripts or impacting running instances.

Exam trap

The trap here is that candidates may confuse stopping instances (which does not delete volumes or snapshots) with deleting resources, or think disabling CloudTrail reduces storage costs, when in fact CloudTrail logs are stored in S3, not EBS, and have separate cost implications.

756
MCQeasy

A startup runs a stateless web tier on Amazon EC2 instances in an Auto Scaling group that spans three Availability Zones. The team wants the application to keep serving requests even if one instance becomes unresponsive, without operator involvement. What should the solutions architect configure?

A.An Application Load Balancer with health checks that route traffic only to healthy targets.
B.A Network Load Balancer configured with a single target group and no health checks.
C.An Elastic IP address associated with each EC2 instance in the Auto Scaling group.
D.An Auto Scaling group with a scheduled scaling policy that adds instances every morning.
AnswerA

The load balancer performs health checks against each target and stops sending traffic to targets that fail, so an unresponsive instance is removed from rotation automatically. Combined with the Auto Scaling group, replacement capacity is launched, keeping the stateless web tier available without human action.

Why this answer

Health-checked load balancing is the mechanism that detects an unresponsive instance and removes it from the traffic path automatically. Because the tier is stateless and already spans three Availability Zones, the Application Load Balancer can shift requests to remaining healthy instances while the Auto Scaling group replaces the failed one.

Exam trap

The trap here is expecting an Elastic IP or scheduled scaling to provide automatic failover for an unhealthy instance, when only health-checked load balancing removes it from service.

757
MCQmedium

A web API runs on an Auto Scaling group (ASG) behind an Application Load Balancer (ALB). During traffic spikes, users experience request timeouts even though CPU stays below 40%. After investigation, you find the ASG often has too few healthy targets to handle the current request rate. Which change will best improve responsiveness during spikes?

A.Keep the ASG scaling policy based on CPU utilization, but increase the ASG min capacity by 50%.
B.Create a target tracking scaling policy using an ALB metric such as RequestCountPerTarget or TargetResponseTime.
C.Enable EC2 detailed monitoring for one-minute granularity and keep CPU scaling.
D.Switch to scaling based on the ASG network out bytes metric only, ignoring ALB response metrics.
AnswerB

Target tracking with an ALB performance metric scales based on the same layer where the problem is observed (requests/latency through the ALB). As traffic spikes, RequestCountPerTarget and/or TargetResponseTime increase; the scaling policy then increases the ASG desired capacity so the ALB has more healthy targets to distribute requests to. That reduces queuing/latency and helps prevent timeouts without waiting for CPU to rise.

Why this answer

The issue is that the ASG has too few healthy targets to handle the request rate, even though CPU is low. A target tracking scaling policy based on RequestCountPerTarget or TargetResponseTime directly aligns scaling with the ALB's view of demand, ensuring the ASG adds instances when request rates spike, regardless of CPU utilization. This addresses the root cause—insufficient capacity to serve incoming requests—rather than relying on a metric (CPU) that does not reflect the bottleneck.

Exam trap

The trap here is that candidates assume CPU utilization is always the best scaling metric, but AWS explicitly tests that ALB-level metrics (RequestCountPerTarget, TargetResponseTime) are more appropriate when the bottleneck is request throughput rather than compute load.

How to eliminate wrong answers

Option A is wrong because increasing the ASG min capacity only raises the baseline number of instances, but does not make the scaling policy responsive to traffic spikes; the ASG will still scale based on CPU, which remains low, so it will not add instances during spikes. Option C is wrong because enabling detailed monitoring (1-minute granularity) improves the frequency of metric data but does not change the fact that CPU utilization is not the correct metric to trigger scaling for this request-rate bottleneck. Option D is wrong because switching to scaling based solely on ASG network out bytes ignores the ALB's request-level metrics, which are more directly correlated with the user-observed timeouts and healthy-target deficit.

758
MCQeasy

A startup runs a customer-facing web application on a single Amazon EC2 instance in one Availability Zone, with the database on the same instance. The founders want the application to survive the failure of that Availability Zone with minimal changes and no server management for the database tier. Which action should the solutions architect take first?

A.Create an Amazon Machine Image of the instance and copy it to a second Region for disaster recovery.
B.Enable termination protection on the EC2 instance and take regular Amazon EBS snapshots.
C.Place an Application Load Balancer in front of the single instance and enable sticky sessions.
D.Move the database to Amazon RDS with a Multi-AZ DB instance deployment and place the application instances in an Auto Scaling group spanning multiple Availability Zones.
AnswerD

Moving the database to RDS Multi-AZ provides automatic failover to a standby in another Availability Zone and removes server management for the database. Placing the application in a multi-AZ Auto Scaling group ensures compute capacity survives a zonal failure, directly addressing the single-AZ risk with a well-understood pattern.

Why this answer

The core problem is that both the application and database reside in a single Availability Zone. Relocating the database to Amazon RDS with Multi-AZ deployment gives automatic failover and removes database server management, while running the application in a multi-AZ Auto Scaling group provides compute redundancy. Together these changes let the application survive a zonal failure with minimal rework.

Exam trap

The trap here is believing that snapshots or termination protection provide high availability, when they only aid recovery and do not keep the workload running through an Availability Zone failure.

759
MCQeasy

A startup runs a web application on Amazon EC2 instances behind an Application Load Balancer. The security team wants to encrypt data in transit between clients and the load balancer using a certificate managed by AWS, with minimal operational overhead. Which solution meets these requirements?

A.Import a self-signed certificate into AWS Certificate Manager (ACM) and attach it to the load balancer's HTTPS listener.
B.Request a public certificate in AWS Certificate Manager (ACM), validate it via DNS, and attach it to an HTTPS listener on the Application Load Balancer.
C.Use an Amazon-issued certificate from IAM and attach it to the load balancer's HTTPS listener.
D.Generate a certificate using AWS KMS, store it in AWS Secrets Manager, and configure the load balancer to retrieve it at runtime.
AnswerB

ACM provides free public certificates that are automatically renewed when validated via DNS and used with integrated services like ALB. Attaching the certificate to an HTTPS listener encrypts client-to-load-balancer traffic. This is the lowest-overhead, fully managed approach that meets the encryption requirement.

Why this answer

AWS Certificate Manager (ACM) issues free public certificates that can be automatically renewed when DNS validation is used, and it integrates directly with Application Load Balancers. Attaching an ACM certificate to an HTTPS listener encrypts client traffic with minimal operational effort, satisfying the requirement.

Exam trap

The trap here is assuming that imported or self-signed certificates in ACM are automatically renewed, when ACM only auto-renews certificates it issued and that use DNS validation.

760
MCQmedium

A media company uploads raw video thumbnails to an S3 bucket every hour. The application needs these thumbnails for active browsing for the first 7 days. After day 7, access becomes rare. Requirements: - Objects must remain available in S3 for at least 180 days total. - After day 7, the team can tolerate retrieval latency in the range of minutes to hours. - They want to minimize storage cost while keeping the ability to read objects (no application changes required). Which storage strategy is the most cost-optimized fit?

A.Use a bucket-level lifecycle rule to transition objects to S3 Standard-IA on day 7 and then expire them after day 180.
B.Use a lifecycle rule to transition objects to S3 Glacier Flexible Retrieval after day 7 and expire them after day 180.
C.Keep all objects in S3 Standard for 180 days, and enable S3 Intelligent-Tiering only if the bucket’s access frequency is above a threshold.
D.Use a lifecycle rule to transition objects to S3 Glacier Instant Retrieval after day 7 and expire them after day 180.
AnswerB

Glacier Flexible Retrieval is designed for infrequent access and supports restore times compatible with minutes to hours. Transitioning after day 7 reduces storage cost for the long period where access is rare, while expiring at day 180 satisfies the 180-day retention requirement. The application can still use S3 GetObject; retrieval simply takes longer due to the archival tier.

Why this answer

S3 Glacier Flexible Retrieval provides retrieval times from minutes to hours, which matches the tolerance for rare access after day 7, and offers the lowest storage cost among the options for data that is rarely accessed. A lifecycle rule transitions objects from S3 Standard (used for the first 7 days of active browsing) to Glacier Flexible Retrieval on day 7, then expires them after day 180, meeting the 180-day retention requirement without application changes.

Exam trap

The trap here is that candidates often choose S3 Glacier Instant Retrieval (Option D) because of the word 'Instant,' overlooking that the requirement explicitly tolerates minutes-to-hours latency, making the cheaper Glacier Flexible Retrieval the better cost-optimized choice.

How to eliminate wrong answers

Option A is wrong because S3 Standard-IA is designed for infrequent access but still incurs higher storage costs than Glacier Flexible Retrieval for data that is accessed rarely (minutes-to-hours latency is acceptable), and it does not provide the lowest cost for this use case. Option C is wrong because keeping all objects in S3 Standard for 180 days is significantly more expensive than transitioning to a colder storage class, and S3 Intelligent-Tiering is not cost-optimized for a predictable access pattern (active for 7 days, then rarely accessed) as it adds monitoring costs and may not move objects to the cheapest tier quickly enough. Option D is wrong because S3 Glacier Instant Retrieval is designed for millisecond retrieval, which is unnecessary and more expensive than Glacier Flexible Retrieval when minutes-to-hours latency is acceptable, thus not the most cost-optimized choice.

761
MCQmedium

Based on the exhibit, the payment worker sometimes processes the same SQS Standard message more than once after a timeout. What change best prevents duplicate charges while keeping the queue architecture?

A.Increase the SQS visibility timeout to 15 minutes and leave the worker unchanged.
B.Replace the Standard queue with a FIFO queue and rely only on message ordering.
C.Make the payment workflow idempotent by recording a unique order key before charging.
D.Add a second consumer so duplicate messages are processed faster.
AnswerC

SQS Standard queues are at-least-once delivery, so duplicate messages are always possible. The correct safeguard is idempotency: store a unique order or payment request key, check whether that key has already been processed, and only perform the charge the first time it is seen. Any later delivery is safely ignored.

Why this answer

Making the payment workflow idempotent ensures that even if the same SQS Standard message is processed more than once (due to a visibility timeout), the duplicate charge is prevented by checking a unique order key before processing. This is the most robust solution for handling at-least-once delivery semantics of Standard queues without changing the queue architecture.

Exam trap

The trap here is that candidates often think increasing the visibility timeout (Option A) or switching to a FIFO queue (Option B) will solve duplicate processing, but they overlook that the root cause is the worker's timeout behavior, which requires application-level idempotency to prevent duplicate charges.

Why the other options are wrong

A

Increasing the visibility timeout to 15 minutes does not prevent duplicate processing; it only reduces the likelihood of timeouts causing duplicates. The worker can still process the same message twice if the timeout expires after 15 minutes, and the change does not address the root cause of duplicate charges.

B

FIFO queues guarantee exactly-once processing and ordering, but the question asks to prevent duplicate charges while keeping the queue architecture. Replacing Standard with FIFO changes the queue type, which may not be desired, and FIFO alone does not prevent duplicate charges if the worker is not idempotent.

D

Adding a second consumer does not prevent duplicate processing; it may even increase the chance of duplicates if both consumers process the same message after a timeout.

When would these options actually be correct?

A

This option would be correct in a scenario where messages are consistently being processed before the visibility timeout expires, but the timeout is too short for the current processing time, causing unnecessary retries. Increasing the timeout to match the maximum processing time would prevent those retries.

B

If the question required ensuring exactly-once processing and message ordering without changing the worker logic, and the architecture allowed replacing the queue type, then using a FIFO queue would be correct. For example: 'A financial application must process payments in strict order without duplicates. Which queue type ensures this?'

D

When the goal is to increase throughput and reduce latency for processing messages, and duplicate processing is acceptable or handled elsewhere (e.g., idempotent workers).

Why candidates pick the wrong answer

A

Candidates may think that a longer visibility timeout guarantees that a message won't be reprocessed, overlooking that timeouts can still occur after the increased duration, and that the fundamental issue of duplicate processing remains unaddressed.

B

Candidates know FIFO queues prevent duplicates and assume that solves the problem, overlooking that the worker must also be idempotent to handle retries, and that the question explicitly says 'keep the queue architecture' (i.e., Standard queue).

D

Candidates may think more consumers reduce the chance of any single message being processed twice, but they overlook that duplicate processing stems from visibility timeout issues, not consumer count.

762
Multi-Selecthard

Multiple teams share one AWS Organization. Finance wants chargeback by project, alerts before overspend, and monthly views by account without manually opening each account. Which three actions best fit? Select three.

Select 3 answers
A.Enforce cost allocation tags on resources and activate them for billing reports.
B.Use AWS Budgets to create alerts and budget actions for each project.
C.Use Cost Explorer or Cost and Usage Reports to analyze spend by account, tag, and service.
D.Put every team in a separate AWS account and ignore tagging.
E.Use CloudTrail trails to estimate spend by resource because it records API calls.
AnswersA, B, C

Enforcing cost allocation tags and activating them for billing reports creates the project dimension Finance needs for chargeback, since AWS only surfaces tag-level costs once tags are activated in Billing. This directly satisfies the stem's requirement to attribute shared-account spend by project rather than by account alone.

Why this answer

Option A is correct because enforcing and activating cost allocation tags (e.g., project tags) is the prerequisite for attributing AWS charges to projects in billing data, enabling accurate chargeback. Option B is correct because AWS Budgets supports cost budgets scoped by tag, account, or service, and can trigger SNS alerts and budget actions (such as applying SCPs or stopping instances) before overspend occurs. Option C is correct because Cost Explorer and Cost and Usage Reports (CUR) provide the consolidated, multi-account analysis by account, tag, and service that Finance needs without manually opening each account, especially when integrated with AWS Organizations.

Option D is not appropriate because merely isolating teams into separate accounts without tagging does not provide project-level chargeback or the required tag-based views. Option E is incorrect because CloudTrail records API activity for auditing, not resource costs, and cannot estimate spend by resource.

Exam trap

The trap here is that candidates may confuse CloudTrail (which records API calls) with AWS Cost Explorer or CUR (which provide actual cost data), leading them to incorrectly select option E for cost estimation.

Why the other options are wrong

D

Putting teams in separate accounts without tagging prevents chargeback by project and requires manual account access for monthly views, failing to meet the requirements for cost allocation and automated reporting.

E

CloudTrail records API calls for auditing, not cost allocation. It does not provide cost or usage data by resource, tag, or project, so it cannot support chargeback, alerts, or monthly views by account.

When would these options actually be correct?

D

If the question asked for the best way to ensure security isolation and prevent resource sharing between teams, with cost tracking done via consolidated billing reports per account, then separate accounts without tagging would be correct.

E

If a question asked for a service to track API activity for security auditing or to identify which user created a resource, CloudTrail would be the correct answer. For example: 'Which service records API calls for operational and risk auditing?'

Why candidates pick the wrong answer

D

Candidates may think separate accounts inherently solve cost tracking, overlooking that chargeback by project still requires tags or other mechanisms to break down costs within an account.

E

Candidates may think CloudTrail can estimate costs because it logs resource creation events, but it lacks pricing data and cannot aggregate spend by tag or account.

763
MCQhard

A healthcare company uses AWS Lambda functions to process sensitive patient data. The functions need to access an Amazon RDS for MySQL database. The security team requires that database credentials are never stored in the Lambda function code or environment variables, and that credentials are automatically rotated every 90 days. The company also wants to minimize the operational overhead of managing the rotation. Which solution should a solutions architect recommend?

A.Store the credentials in AWS Systems Manager Parameter Store as a SecureString parameter and manually rotate them every 90 days.
B.Store the credentials in AWS Secrets Manager and configure automatic rotation using a Lambda rotation function. Grant the Lambda execution role permission to retrieve the secret.
C.Store the credentials in an encrypted Amazon S3 object and have the Lambda function download and decrypt the object at runtime.
D.Use AWS IAM database authentication for RDS and generate temporary credentials using AWS Security Token Service (STS).
AnswerB

AWS Secrets Manager provides built-in support for automatic rotation of database credentials using a Lambda rotation function. It eliminates the need to store credentials in code or environment variables. The Lambda execution role can be granted permission to retrieve the secret at runtime, ensuring secure and automated access with minimal operational overhead.

Why this answer

AWS Secrets Manager is designed for secure storage and automatic rotation of secrets, including database credentials. It integrates with RDS to rotate credentials using a Lambda function, reducing operational overhead. The Lambda function can retrieve the secret at runtime using its execution role, ensuring credentials are never stored in code or environment variables.

Exam trap

The trap here is confusing IAM database authentication with automatic credential rotation; IAM auth uses temporary tokens but does not rotate stored credentials.

764
MCQmedium

Developers for a customer analytics portal need temporary elevated access to production resources for troubleshooting. The security team wants approvals, expiry, and audit logging. Which approach is best?

A.Disable CloudTrail during troubleshooting
B.Attach AdministratorAccess permanently to every developer role
C.Use IAM Identity Center permission sets with time-bound access processes and CloudTrail auditing
D.Create shared administrator access keys for the team
AnswerC

IAM Identity Center permission sets allow you to assign role-based permissions (e.g., AdministratorAccess) to users or groups for a defined session duration, often combined with temporary elevation workflows that require approval and expire automatically. By coupling these time-bound assignments with CloudTrail auditing, every privileged action is attributable to a specific federated user and can be reviewed for anomalies. This reduces standing privilege, supports just-in-time access, and gives the team the temporary administrative power they need without permanently widening the security posture.

Why this answer

AWS IAM Identity Center (formerly AWS SSO) allows you to define permission sets that grant temporary, time-bound elevated access to production resources. Combined with AWS CloudTrail, every access attempt is logged for audit, meeting the security team's requirements for approvals, expiry, and audit logging. This approach follows the principle of least privilege and ensures that elevated permissions are not permanent.

Exam trap

The trap here is that candidates often confuse IAM users with IAM Identity Center, or think that simply enabling CloudTrail (without a proper access control mechanism) is sufficient, but the question specifically requires time-bound access and approvals, which only a centralized identity solution like IAM Identity Center provides.

How to eliminate wrong answers

Option A is wrong because disabling CloudTrail during troubleshooting removes all audit logging, directly violating the security team's requirement for audit logging. Option B is wrong because permanently attaching AdministratorAccess to every developer role grants excessive, permanent privileges, violating the principle of least privilege and the requirement for temporary, time-bound access. Option D is wrong because creating shared administrator access keys eliminates individual accountability, breaks audit trails, and violates the security team's need for approvals and expiry.

765
MCQmedium

A production Amazon RDS database has automated backups enabled. At 10:00 UTC, an application deploy accidentally overwrote a subset of rows due to a faulty migration. The issue is detected at 10:45 UTC. The team confirms that the required retention window is still available. Which approach offers the most resilient and least disruptive way to recover the affected data close to the time of the event?

A.Perform a snapshot restore and attach the restored instance, then manually copy only the affected rows back into the current database.
B.Use point-in-time recovery to restore the database to a timestamp just before 10:00 UTC, then swap application connectivity to the recovered instance.
C.Rely on automated backups to roll forward automatically until the data becomes correct.
D.Disable automated backups going forward to prevent future corruption, then reindex the corrupted table.
AnswerB

Point-in-time recovery (PITR) for Amazon RDS uses automated backups and transaction logs to restore the database to any second within the backup retention period, allowing you to target a timestamp just before 10:00 UTC when the corruption occurred. This minimizes data loss to only the changes made in the seconds immediately preceding the incident, far more precise than a full snapshot. After restoring to a new RDS instance, you swap the application connection string (or use Route 53/RDS Proxy) to point to the recovered instance, enabling a clean recovery with minimal disruption and no manual row copying.

Why this answer

Point-in-time recovery (PITR) allows you to restore the RDS instance to any second within the backup retention window, such as just before the faulty migration at 10:00 UTC. This restores a complete, consistent database state, minimizing data loss and avoiding manual row-by-row recovery. Swapping application connectivity to the restored instance is the least disruptive approach, as it avoids complex manual data merging and reduces downtime.

Exam trap

The trap here is that candidates may choose snapshot restore (Option A) thinking it is faster or simpler, but they overlook that PITR provides a more precise, consistent recovery point without manual data extraction and reinsertion.

How to eliminate wrong answers

Option A is wrong because performing a snapshot restore and manually copying affected rows is error-prone, time-consuming, and risks data inconsistency, especially if the affected rows have dependencies. Option C is wrong because automated backups do not 'roll forward' to correct data corruption; they are used for restore operations, not automatic healing. Option D is wrong because disabling automated backups does not recover lost data and actually increases future risk; reindexing does not restore overwritten rows.

766
MCQhard

Based on the exhibit, which storage choice best matches the workload requirements?

A.Use io2 EBS volumes because they provide the highest durable block storage performance.
B.Use instance store NVMe for the temporary processing workspace.
C.Use Amazon EFS for the workspace so the temporary files survive instance replacement.
D.Use S3 as the working directory and read and write the intermediate files directly there.
AnswerB

Instance store fits a high-IOPS scratch workload where data can be lost safely and rebuilt from S3. The benchmark shows extremely low latency and very high random I/O performance, which is ideal for intermediate transcode files. Because the job can be retried from the source object, persistence is not needed on the local workspace.

Why this answer

Instance store NVMe volumes provide temporary, high-performance block storage directly attached to the EC2 host. For a temporary processing workspace where data does not need to persist beyond the instance lifecycle, instance store offers the lowest latency and highest throughput, making it the best match for the workload requirements.

Exam trap

The trap here is that candidates often choose io2 EBS volumes (Option A) because they associate 'highest durable block storage' with 'best performance,' failing to recognize that durability and persistence are unnecessary for temporary data, and that instance store provides superior raw performance for ephemeral workloads.

Why the other options are wrong

A

The workload requires temporary processing workspace, which does not need high durability. io2 EBS volumes are designed for high durability and consistent performance, but they are unnecessary and more expensive for temporary data that can be regenerated.

C

Amazon EFS provides persistent shared storage, but temporary processing workspace files do not need to survive instance replacement; using EFS adds unnecessary cost and latency compared to instance store.

D

Using S3 as a working directory for intermediate files would introduce high latency and cost due to frequent read/write operations, and S3 does not support file locking or low-latency random access required for temporary processing workspaces.

When would these options actually be correct?

A

A question requiring a persistent, high-performance block storage for a critical database with strict durability requirements, where the data must survive instance failures and be replicated across Availability Zones.

C

A question where the workload requires a shared file system accessible from multiple EC2 instances simultaneously, with files needing to persist across instance terminations, such as a content management system or a shared development environment.

D

In a scenario where the workload involves storing large, immutable intermediate files that are accessed infrequently and need to be shared across multiple instances, S3 would be the correct choice for durability and scalability.

Why candidates pick the wrong answer

A

Candidates may associate 'high performance' with io2 volumes and overlook that the workload is temporary and does not require durability, leading to an over-engineered and costly choice.

C

Candidates may think that temporary files should be stored durably to avoid data loss, overlooking that the workload explicitly states the workspace is temporary and does not require persistence.

D

Candidates may think S3's durability and scalability make it suitable for any storage need, overlooking the performance and access pattern requirements of temporary processing workspaces.

767
MCQmedium

A company runs an Amazon Aurora DB cluster with a Multi-AZ deployment. The application is configured with a hard-coded endpoint that points to the current writer *DB instance* (an instance-specific endpoint), rather than the Aurora cluster writer endpoint. During an unexpected AZ failure, Aurora promotes the standby to become the new writer. However, the application continues to fail to connect until an operator updates the hard-coded endpoint. What change most directly improves resiliency so the application automatically reconnects after failover?

A.Keep using the writer DB instance endpoint, but increase the client connection timeout.
B.Connect using the Aurora cluster writer endpoint so DNS resolves to the current writer after failover.
C.Disable Multi-AZ failover and rely on manual snapshot restore to bring the database back online.
D.Enable cross-Region read replicas and route application traffic to the replica during the outage.
AnswerB

Aurora cluster endpoints are designed to provide continuity across failovers. The Aurora cluster writer endpoint (writer endpoint for the cluster) updates so DNS resolves to the promoted writer. The application can reconnect without manual endpoint changes.

Why this answer

The Aurora cluster writer endpoint is a DNS name that always resolves to the current writer instance in the cluster, even after a failover. By using this endpoint instead of a hard-coded instance-specific endpoint, the application automatically reconnects to the new writer without manual intervention, directly improving resiliency.

Exam trap

The trap here is that candidates may confuse the instance-specific endpoint with the cluster writer endpoint, or think that increasing timeouts or using read replicas can solve a writer failover issue, when the core problem is the hard-coded reference to a specific instance that no longer exists.

How to eliminate wrong answers

Option A is wrong because increasing the client connection timeout does not change the fact that the hard-coded endpoint points to a failed instance; the connection will still fail after the timeout expires. Option C is wrong because disabling Multi-AZ failover and relying on manual snapshot restore would cause significant downtime and data loss, directly contradicting the goal of improving resiliency. Option D is wrong because cross-Region read replicas are read-only and cannot accept writes; routing application traffic to a read replica during an outage would not allow the application to write data, and it does not address the failover of the writer instance.

768
Multi-Selecthard

A latency-sensitive video platform uploads large files to S3 from users around the world. Which two features can improve upload performance? The design must avoid adding custom operational scripts.

Select 2 answers
A.S3 Object Lock
B.S3 Transfer Acceleration
C.S3 multipart upload
D.S3 Inventory
AnswersB, C

S3 Transfer Acceleration uses AWS edge locations and a dedicated, optimized network backbone to shorten the data path between your client and the S3 bucket. It routes uploads over the AWS global network instead of the public internet, reducing round-trip times and jitter, especially for cross-continent or high-latency connections. This makes it the right choice for latency-sensitive, large-file uploads over long distances.

Why this answer

S3 Transfer Acceleration (B) uses AWS edge locations to route uploads over optimized network paths, reducing latency for users far from the destination bucket. S3 Multipart Upload (C) allows parallel uploads of file parts, improving throughput and resilience for large files. Both features enhance upload performance without requiring custom scripts.

Exam trap

The trap here is that candidates may confuse S3 Transfer Acceleration with CloudFront’s content delivery features, or assume S3 Object Lock or Inventory could somehow improve upload speed, but neither addresses network latency or throughput for uploads.

769
Multi-Selecthard

A retailer runs a reporting-heavy relational app on Amazon RDS MySQL. Peak dashboard traffic lasts only three hours each day, but the database is sized for the peak all day. The business wants lower cost without rewriting the application. Which three actions are best? Select three.

Select 3 answers
A.Right-size the writer based on actual utilization instead of peak guesses.
B.Add read replicas and direct dashboard traffic away from the writer.
C.Evaluate Aurora MySQL if the current replica-heavy design would be cheaper there.
D.Migrate to DynamoDB immediately because every relational workload is more expensive.
E.Increase provisioned IOPS permanently so the monthly bill drops.
AnswersA, B, C

Right-sizing the writer to actual utilisation rather than peak guesses reduces instance cost while preserving the existing MySQL application. It directly addresses the stem's constraint that the database is over-provisioned for all-day peak, requiring no application rewrite.

Why this answer

Option A is correct because right-sizing the writer instance to match actual CPU, memory, and IOPS utilization rather than peak-day guesses directly reduces the largest cost component of an RDS MySQL deployment without changing the application. Option B is correct because adding read replicas and routing dashboard (read-only) queries to them offloads the writer, letting the writer be smaller and cheaper while still serving peak reporting traffic. Option C is correct because evaluating Aurora MySQL is a valid cost-optimization step: Aurora's distributed storage and replica model can be cheaper for replica-heavy read workloads, and it remains MySQL-compatible so no application rewrite is needed.

Option D is wrong because DynamoDB is a NoSQL service that would require rewriting the relational application, which the business explicitly wants to avoid. Option E is wrong because permanently increasing provisioned IOPS raises, not lowers, the monthly bill.

Exam trap

The trap here is that candidates assume DynamoDB is always cheaper for any workload, ignoring the need for application rewrites and the relational reporting requirements, while also overlooking that increasing IOPS always raises costs rather than lowering them.

770
MCQhard

A genomics company stores about 400 TB of compressed research data in Amazon S3 and runs a nightly analysis job on a fleet of EC2 instances in the same Region. The job reads the entire dataset every night, and the team wants to reduce the time the fleet spends waiting on storage without changing the data format or the S3 bucket. Which change best improves read throughput for the fleet?

A.Configure the S3 requests to use byte-range fetches and issue many concurrent GET requests across the dataset.
B.Change the storage class of the dataset to S3 Glacier Instant Retrieval for faster reads.
C.Mount the bucket with a file gateway and read the dataset as files over NFS.
D.Enable S3 Transfer Acceleration and read the dataset through the accelerated endpoint.
AnswerA

S3 scales throughput with the number of concurrent requests and the spread of key prefixes, and byte-range GETs let each instance fetch different portions in parallel. Issuing many concurrent requests across the dataset uses the aggregate bandwidth available to the fleet, which shortens the time spent waiting on storage without altering the data format or the bucket.

Why this answer

S3 performance scales horizontally with concurrent requests, so a fleet reading 400 TB should issue many parallel, byte-range GETs spread across the key space rather than a few large sequential reads. That pattern uses the aggregate bandwidth of all instances and reduces time spent waiting on storage, while leaving the data format and bucket unchanged as required.

Exam trap

The trap here is assuming a transfer-acceleration or storage-class feature increases read speed for clients already inside the same Region, when the real lever is request concurrency and key distribution.

771
MCQhard

Based on the exhibit, duplicate payment charges occasionally occur when the worker times out after the charge is submitted but before the message is deleted. What change best prevents duplicate charges while keeping retry behavior?

A.Switch the queue to FIFO and rely on content-based deduplication to guarantee exactly-once processing.
B.Make the consumer idempotent by storing a processed payment key and rejecting repeat charges.
C.Reduce the visibility timeout so the message becomes available again sooner after a timeout.
D.Add a dead-letter queue and disable retries so the message is never processed twice.
AnswerB

The worker can still receive the same message more than once because SQS Standard is at-least-once delivery and the delete happened after the charge. Idempotency is the correct safety control because it prevents the payment from being applied twice even when the message is retried. A processed-payment record or conditional write lets retries remain possible without creating duplicate charges.

Why this answer

Making the consumer idempotent ensures that even if the same message is processed more than once (due to a timeout after the charge is submitted but before the message is deleted), the duplicate charge will be rejected. By storing a processed payment key (e.g., a unique transaction ID) and checking it before processing, the system can safely retry without causing duplicate payments. This approach preserves retry behavior while preventing duplicates, which is the core requirement.

Exam trap

The trap here is that candidates often assume FIFO queues with deduplication guarantee exactly-once processing, but they fail to recognize that deduplication only prevents duplicate message delivery, not duplicate processing when the consumer times out after processing but before acknowledging the message.

Why the other options are wrong

A

FIFO queues with content-based deduplication provide exactly-once delivery, but the issue here is a timeout after submission but before deletion, which can still cause duplicate processing if the worker retries. FIFO deduplication does not prevent duplicate charges if the same message is sent again after a timeout, as deduplication is based on message content within a 5-minute window, not on processing state.

C

Reducing the visibility timeout would cause the message to reappear sooner after a timeout, increasing the likelihood of duplicate processing rather than preventing it. It does not address the root cause of duplicate charges when the worker times out after submitting the charge.

D

Adding a dead-letter queue and disabling retries prevents duplicate charges by eliminating retries, but the question explicitly requires keeping retry behavior. This option removes retries, which violates the requirement.

When would these options actually be correct?

A

A question where the requirement is to guarantee exactly-once delivery of messages in a distributed system, and the application can tolerate a 5-minute deduplication window. For example: 'A financial application must ensure that payment requests are processed exactly once, even if the producer sends duplicate messages. Which queue type should be used?'

C

This option would be correct in a scenario where the goal is to minimize latency for retrying failed messages, and duplicate processing is acceptable or handled elsewhere. For example, a question asking 'How to ensure a message is retried quickly after a worker failure?' would make reducing visibility timeout the best answer.

D

A question where the requirement is to prevent duplicate processing of poison-pill messages that cannot be handled successfully, and retries are not needed. For example: 'An order processing system fails repeatedly on certain malformed messages. What is the most cost-effective way to isolate these messages for manual inspection without retrying them?'

Why candidates pick the wrong answer

A

Candidates often associate FIFO queues with exactly-once processing and assume that deduplication solves all duplicate issues, without considering that the duplicate here arises from a retry after a timeout, not from duplicate message sends.

C

Candidates may think that making the message available again faster will reduce the chance of duplicate charges by allowing the worker to delete the message before a new attempt, but they overlook that the duplicate charge has already occurred before the timeout.

D

Candidates may think that a dead-letter queue is a standard solution for preventing duplicates, and disabling retries seems like a direct way to avoid reprocessing, overlooking the explicit requirement to retain retry behavior.

772
Matchingmedium

A team wants a web application to keep serving traffic if one Availability Zone fails. Match each architecture element to the resilience behavior it provides.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Stop sending requests to unhealthy targets and keep only healthy instances in rotation.

Launch replacement instances in healthy AZs when capacity is lost.

Maintain a synchronous standby in another AZ and fail over automatically.

Allow instances to be replaced without losing user sessions that are stored elsewhere.

Why these pairings

These pairs match architecture elements with their resilience behaviors for surviving an Availability Zone failure, focusing on AWS services that provide high availability and fault tolerance.

773
MCQhard

A data engineering team ingests a continuous stream of clickstream events into Amazon Kinesis Data Streams. Downstream consumers process the events, but the team observes that a single consumer is handling a disproportionate share of the records, causing hot shards and throttling. The team wants the stream to distribute records as evenly as possible across shards. Which change should the team make?

A.Increase the retention period of the stream to 365 days to smooth out the load.
B.Enable enhanced fan-out on the stream so each consumer gets dedicated throughput.
C.Switch the producer to use an explicit partition key equal to the event's source IP address.
D.Increase the number of open shards and configure the producer to use a random partition key for each record.
AnswerD

Record distribution across shards is determined by the partition key: Kinesis hashes the key and maps it to a shard. A random or highly varied partition key spreads records across all shards, avoiding a hot shard, and adding shards increases the total capacity so the workload can be absorbed evenly.

Why this answer

Skew in Kinesis comes from the partition key mapping many records to the same shard hash range. Choosing a high-cardinality, evenly distributed key such as a random value spreads records across all shards, and adding shards provides more capacity. Together these changes balance the stream and relieve hot-shard throttling.

Exam trap

The trap here is blaming the consumers and reaching for enhanced fan-out when the imbalance originates in how the producer chooses partition keys.

774
MCQmedium

A fintech company runs a containerized payment API on Amazon ECS with AWS Fargate. The security team requires that the API access a stored database credential without hardcoding it in the task definition or environment variables. The credential must be encrypted at rest and automatically rotated every 90 days. The API also needs to retrieve the credential at container startup with minimal latency. Which solution meets these requirements?

A.Store the credential in an encrypted Amazon S3 object and grant the ECS task role s3:GetObject. Configure a Lambda function to rotate the credential every 90 days.
B.Store the credential in AWS Systems Manager Parameter Store as a SecureString, and grant the ECS task role ssm:GetParameter. Enable automatic rotation through Parameter Store.
C.Store the credential in an encrypted Amazon EBS volume attached to the Fargate task, and grant the task role permission to mount the volume.
D.Store the credential in AWS Secrets Manager, enable automatic rotation, and grant the ECS task role permission to call secretsmanager:GetSecretValue.
AnswerD

AWS Secrets Manager supports native automatic rotation for supported databases and can rotate credentials every 90 days. The ECS task role can be granted least-privilege access to GetSecretValue, and the container retrieves the secret at startup. Secrets Manager encrypts secrets at rest using KMS and integrates with IAM, satisfying the encryption and access-control requirements without hardcoding credentials.

Why this answer

AWS Secrets Manager is designed for storing, encrypting, and automatically rotating secrets such as database credentials. By using the ECS task role to call GetSecretValue, the container retrieves the credential at startup without embedding it in the task definition. Native rotation every 90 days satisfies the compliance requirement, and KMS encryption at rest plus IAM policies enforce least privilege.

Exam trap

The trap here is assuming Parameter Store SecureString provides automatic rotation, when rotation must be custom-built with Lambda and EventBridge.

775
MCQmedium

A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The team wants the control to be enforceable during normal operations.

A.Aurora Global Database
B.A single-AZ Aurora cluster
C.An ElastiCache Redis replica
D.Manual snapshots copied monthly
AnswerA

Aurora Global Database replicates data at the storage layer to secondary Regions with typical latency under one second, using an asynchronous but dedicated replication channel. It supports both planned switchover and unplanned failover promotion, achieving RPO of seconds and RTO of minutes, which far exceeds the ticket system's need for fast failover. Secondary Regions can also serve local reads, improving both availability and recovery performance.

Why this answer

Aurora Global Database is designed for cross-Region disaster recovery with a typical RPO of 1 second or less, using storage-based replication that does not impact database performance. It provides fast failover to a secondary Region and allows the primary Region to enforce write control during normal operations, meeting the low RPO and enforceable control requirements.

Exam trap

The trap here is that candidates may confuse Aurora Global Database with cross-Region read replicas or manual snapshot copy strategies, underestimating the RPO and failover speed requirements for disaster recovery.

How to eliminate wrong answers

Option B is wrong because a single-AZ Aurora cluster lacks any cross-Region replication or failover capability, resulting in no DR protection and an RPO that depends on manual backups. Option C is wrong because ElastiCache Redis is an in-memory cache, not a persistent database, and its cross-Region replication (Global Datastore) does not provide the same transactional consistency or DR guarantees as Aurora Global Database for a ticket booking system. Option D is wrong because manual snapshots copied monthly would yield an RPO of up to 30 days, far exceeding the low RPO requirement, and they require manual intervention for recovery, which is not fast.

776
MCQeasy

A team stores application logs in Amazon CloudWatch Logs. They enabled long retention and detailed dashboards, resulting in higher-than-expected monthly spend. Compliance requires retaining logs for 90 days, but operations only needs aggregated views. Which change most directly reduces CloudWatch Logs cost while meeting the requirement?

A.Set the CloudWatch Logs log group retention period to 90 days for the relevant log groups.
B.Disable VPC flow logs so the applications stop producing logs automatically.
C.Increase the logging level to DEBUG to reduce the number of log events by batching them.
D.Turn off CloudWatch alarms so logs stop being ingested into CloudWatch Logs.
AnswerA

CloudWatch Logs storage charges are calculated per GB per month, and log groups default to never expiring unless you configure a retention policy. Setting the retention period to 90 days on the specific log groups the applications write to causes CloudWatch Logs to automatically delete log events older than 90 days, which progressively reduces the stored volume and therefore the monthly storage cost while still preserving the required 90 days of logs. This directly addresses the storage cost driver: how long data is retained.

Why this answer

Setting the CloudWatch Logs log group retention period to 90 days directly reduces storage costs by automatically expiring logs after the compliance-required duration. This eliminates the cost of storing logs beyond 90 days, which was the primary driver of the higher-than-expected spend, while still retaining the data for the mandated period and allowing aggregated views via dashboards.

Exam trap

The trap here is that candidates may confuse log retention settings with log ingestion controls, mistakenly thinking that disabling alarms or changing log levels will reduce costs, when in fact the most direct and compliant cost-saving measure is to adjust the retention period.

How to eliminate wrong answers

Option B is wrong because disabling VPC Flow Logs stops the production of network-level logs, but the question states the team stores 'application logs' in CloudWatch Logs, not VPC Flow Logs; this action would not address the cost of existing application log ingestion and retention, and it would break compliance if those logs are required. Option C is wrong because increasing the logging level to DEBUG actually generates more log events per operation, not fewer, and batching does not reduce the number of events; it would increase costs due to higher ingestion volume. Option D is wrong because turning off CloudWatch alarms does not stop log ingestion; alarms are separate from log data ingestion and retention, so logs would continue to be ingested and stored, incurring the same costs.

777
MCQhard

A financial services company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The security team wants to inspect incoming requests for common web exploits and block malicious traffic before it reaches the application. They also need to monitor for SQL injection attempts and receive near-real-time metrics. Which AWS service should be used to meet these requirements?

A.AWS WAF
B.AWS Network Firewall
C.AWS Shield Advanced
D.Amazon GuardDuty
AnswerA

AWS WAF inspects HTTP(S) requests and can block common web exploits using managed rule groups, including rules for SQL injection. It integrates with Application Load Balancer, CloudFront, and API Gateway, and publishes metrics to CloudWatch for near-real-time monitoring. This directly addresses the requirement to inspect and block malicious traffic before it reaches the application.

Why this answer

The need to inspect HTTP requests for SQL injection and common web exploits, block them inline, and emit metrics points to AWS WAF. It attaches to the Application Load Balancer and evaluates requests against rule groups, including managed rules for SQL injection and other OWASP threats. Shield Advanced, GuardDuty, and Network Firewall address DDoS, threat detection, and network-layer filtering, respectively, not application-layer request inspection.

Exam trap

The trap here is confusing DDoS protection or threat detection with inline application-layer request filtering.

778
MCQeasy

You have an EC2 instance in private subnets with no NAT Gateway. The instance must access an Amazon S3 bucket (for example, to read configuration files) without sending traffic to the public internet. What VPC endpoint type should you use for S3?

A.Create a Gateway VPC endpoint for the S3 service
B.Create an Interface VPC endpoint (powered by PrivateLink) for S3
C.Use a Transit Gateway to route to S3 over the internet
D.Place a NAT Gateway and restrict security group egress to port 443 to reduce exposure
AnswerA

S3 uses a Gateway VPC endpoint type. Gateway endpoints integrate with your VPC route tables so that traffic destined for S3 is routed privately within the VPC, avoiding the need for NAT Gateway and public internet egress for S3 access.

Why this answer

A Gateway VPC endpoint is the correct choice because it allows EC2 instances in a private subnet to access S3 without traversing the public internet. It uses prefix lists and route table entries to direct S3 traffic through AWS's internal network, and it does not require a NAT gateway, internet gateway, or public IP addresses. Gateway endpoints are free of charge and scale automatically, making them ideal for private subnet access to S3 and DynamoDB.

Exam trap

The trap here is that candidates often confuse Gateway VPC endpoints with Interface VPC endpoints, assuming S3 requires a private IP address like other AWS services, but S3 and DynamoDB are the only services that support Gateway endpoints, which are simpler and free.

Why the other options are wrong

B

Interface VPC endpoints are used for services that require PrivateLink, but S3 supports Gateway endpoints which are free and route traffic within the AWS network without traversing the internet. Interface endpoints incur hourly charges and data processing fees, making them unnecessary for S3.

C

A Transit Gateway routes traffic between VPCs and on-premises networks, not to S3 over the internet. It does not provide private connectivity to S3; using it to route over the internet would still require a NAT Gateway or Internet Gateway, which violates the requirement of no public internet traffic.

D

A NAT Gateway would route traffic to the public internet, which violates the requirement to avoid sending traffic to the public internet. Additionally, the instance is in a private subnet with no NAT Gateway, so adding one would incur cost and complexity without meeting the requirement.

When would these options actually be correct?

B

When accessing an S3 bucket from on-premises via AWS Direct Connect or VPN, or when you need to access S3 from another AWS region or from a VPC in a different account using PrivateLink. Also correct if you require fine-grained security group controls for S3 access.

C

A Transit Gateway would be correct in a scenario where you need to connect multiple VPCs and on-premises networks to a central inspection VPC, or to route traffic between them privately. For example, when interconnecting VPCs across accounts or regions for hybrid networking.

D

This option would be correct in a scenario where an EC2 instance in a private subnet needs to access the internet (e.g., for software updates or API calls) and you want to control outbound traffic using security groups. The question would specify that internet access is required and cost is not a primary concern.

Why candidates pick the wrong answer

B

Candidates may confuse Interface endpoints (powered by PrivateLink) as the default or more modern option, not realizing that S3 and DynamoDB uniquely support Gateway endpoints, which are simpler and cheaper for VPC-to-S3 access.

C

Candidates may think Transit Gateway can route traffic to any destination, including S3, and might confuse it with a VPC endpoint. They may also assume that Transit Gateway inherently provides private connectivity to AWS services without understanding its actual use case.

D

Candidates may think a NAT Gateway is necessary for any outbound traffic from private subnets, or they may confuse the need for internet access with the requirement to keep traffic private. They might also assume that S3 access always requires internet routing.

779
MCQeasy

A backend API uses an AWS Lambda function behind API Gateway. The first requests after every weekly deployment experience cold starts, causing p95 latency spikes for a few minutes. Which configuration most directly prevents those cold starts for the published version?

A.Increase the Lambda memory size only, without changing how Lambda is invoked
B.Use Lambda provisioned concurrency for the version via an alias
C.Enable dead-letter queues (DLQ) to retry failed cold starts
D.Attach a CloudFront distribution to cache API Gateway responses for 5 minutes
AnswerB

Provisioned concurrency keeps Lambda execution environments initialized and ready for a specific published version. By attaching it to an alias (for example, pointing the alias used by API Gateway to the new version), you pre-warm environments so the first requests after deployment are served without cold-start initialization.

Why this answer

Provisioned concurrency initializes a specified number of Lambda execution environments ahead of time, so that when the published version is invoked via an alias, there are no cold starts. This directly addresses the latency spikes caused by cold starts after a deployment, as the function is kept warm and ready to handle requests immediately.

Exam trap

The trap here is that candidates may confuse provisioned concurrency with reserved concurrency, which only limits the maximum number of concurrent executions but does not prevent cold starts.

Why the other options are wrong

A

Increasing memory size reduces cold start duration but does not prevent cold starts from occurring; it only makes them shorter. The question asks to 'prevent' cold starts, not mitigate their impact.

C

Dead-letter queues (DLQ) are used to capture events that fail processing after multiple retries, not to prevent cold starts. Cold starts occur due to initialization latency, not invocation failures, so DLQs do not address the latency spike.

D

CloudFront caching reduces latency for repeated requests by serving cached responses, but it does not prevent cold starts for the Lambda function. Cold starts occur when Lambda initializes a new execution environment, which happens regardless of API Gateway or CloudFront caching.

When would these options actually be correct?

A

A question asks: 'A Lambda function processes user uploads and experiences high latency during cold starts. Which configuration reduces the duration of cold starts without changing the invocation pattern?' In that case, increasing memory would be correct.

C

A correct scenario: A Lambda function processes messages from an SQS queue, and some messages fail due to transient errors. Enabling a DLQ would capture those failed messages after the retry limit is exhausted, allowing later analysis or reprocessing.

D

A question where the API returns static or slowly-changing data and the goal is to reduce latency and API Gateway load for repeated requests. For example: 'A web application serves a static JSON configuration file via API Gateway. How can you reduce latency for users and decrease the number of requests reaching the backend?'

Why candidates pick the wrong answer

A

Candidates know that more memory reduces cold start latency, so they assume it prevents cold starts entirely, overlooking that provisioned concurrency is needed to keep functions warm.

C

Candidates may think DLQs can 'retry' cold starts, misunderstanding that cold starts are not errors but initialization delays, and that DLQs handle failed invocations, not latency issues.

D

Candidates may think that caching responses will mask the cold start latency by serving cached data during the cold start period, but cold starts affect the first request after deployment, which is not cached yet.

780
MCQmedium

A site serves static assets (JS/CSS) through CloudFront from an S3 origin. After a recent frontend change, CloudFront shows a cache hit ratio below 20%. In CloudFront access logs, requests to the same asset URL path differ by a query parameter named rnd (a random value appended by the app on every request). The origin content is identical regardless of rnd. What is the best CloudFront configuration change to restore effective caching?

A.Increase the origin response Cache-Control max-age header on S3 so CloudFront caches longer even with different rnd values.
B.Create a custom CloudFront Cache Policy that does not include the rnd query parameter in the cache key (whitelist only required parameters, or forward no query strings).
C.Disable compression on CloudFront so the response body is identical byte-for-byte and cache hits improve.
D.Switch the origin from S3 to an ALB so CloudFront can cache based on ALB target health checks instead of the query string.
AnswerB

CloudFront caching effectiveness depends on the cache key. Since rnd does not change the content returned by the S3 origin, excluding rnd from the cache key allows many requests for the “same” asset to map to the same cached object. This removes cache fragmentation and restores a higher hit ratio without changing application content correctness.

Why this answer

The rnd query parameter makes each request appear unique to CloudFront, causing a cache miss for every request even though the underlying content is identical. By creating a custom cache policy that either forwards no query strings or whitelists only required parameters, CloudFront will ignore the rnd parameter when computing the cache key, allowing it to serve cached responses and dramatically improve the cache hit ratio.

Exam trap

The trap here is that candidates often think increasing cache duration (Option A) or disabling compression (Option C) will fix cache misses, when the real issue is that the query parameter is being included in the cache key, making every request unique.

How to eliminate wrong answers

Option A is wrong because increasing Cache-Control max-age only tells the browser and edge how long to cache the response, but it does not change the cache key; CloudFront still treats URLs with different rnd values as distinct objects, so each request will be a cache miss. Option C is wrong because disabling compression does not affect the cache key; CloudFront already caches compressed and uncompressed versions separately based on the Accept-Encoding header, and the issue here is the query string, not compression. Option D is wrong because switching to an ALB does not solve the query-string-based cache key problem; CloudFront would still see different rnd values as different cache keys, and ALBs are not designed to improve CloudFront caching behavior.

781
Multi-Selectmedium

A logistics company runs an order-tracking service on Amazon EC2 instances that write state to an Amazon DynamoDB table. A recent incident showed that a single Availability Zone failure caused the service to lose capacity, and the team also discovered that a developer accidentally deleted a production table. The architect must improve both Availability Zone resilience and protection against accidental table deletion. (Choose two.)

Select 2 answers
A.Enable DynamoDB point-in-time recovery on the table and attach a resource-based policy that denies the dynamodb:DeleteTable action to non-administrative principals.
B.Create a DynamoDB global secondary index on the partition key used by the tracking queries and project all attributes into the index.
C.Enable DynamoDB Streams on the table and write a consumer that copies every change into an Amazon S3 bucket for long-term retention.
D.Deploy the EC2 instances in an Auto Scaling group that spans multiple Availability Zones and attach the instances to an Application Load Balancer.
E.Convert the DynamoDB table to use provisioned capacity mode with auto scaling so that read and write capacity automatically adjusts during traffic spikes.
AnswersA, D

Point-in-time recovery allows restoration of the table to any second within the previous 35 days, which recovers data after an accidental deletion. Adding an IAM policy that denies dynamodb:DeleteTable to ordinary principals reduces the chance of the same mistake recurring, so together they address the accidental-deletion risk.

Why this answer

Zone resilience for the compute tier comes from running instances across multiple Availability Zones behind a load balancer, so a single zone loss does not remove all capacity. Accidental table deletion is mitigated by enabling point-in-time recovery, which allows restore to a recent point in time, and by restricting who holds the delete-table permission so the mistake is far less likely to recur.

Exam trap

The trap here is assuming DynamoDB needs multi-AZ configuration like a relational database, when DynamoDB already replicates data across zones and the real gaps are compute placement and deletion protection.

782
MCQhard

A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The team wants the control to be enforceable during normal operations.

A.Amazon EFS with mount targets in multiple Availability Zones
B.S3 mounted as a POSIX file system without a file gateway
C.Instance store volumes
D.An EBS volume attached to all instances
AnswerA

EFS is regional file storage and supports mount targets across AZs.

Why this answer

Amazon EFS provides a fully managed, POSIX-compliant NFS file system that can be mounted concurrently on multiple Linux EC2 instances across different Availability Zones. By creating mount targets in each AZ, the file system remains accessible even if one AZ fails, because the other mount targets continue to serve traffic. EFS also supports lifecycle policies and IAM enforcement to control access during normal operations, meeting the requirement for enforceable control.

Exam trap

The trap here is that candidates often confuse EBS multi-attach (which is limited to specific instance types and a single AZ) with the cross-AZ shared file system capability that only EFS provides, or they mistakenly think S3 with a FUSE mount is a reliable POSIX file system for production workloads.

How to eliminate wrong answers

Option B is wrong because mounting S3 as a POSIX file system (e.g., using s3fs-fuse) does not provide true POSIX semantics (e.g., no file locking, eventual consistency) and is not designed for shared file storage across AZs with high availability during an AZ failure. Option C is wrong because instance store volumes are ephemeral, tied to a single EC2 instance, and data is lost if the instance stops or fails; they cannot be shared across instances or survive an AZ failure. Option D is wrong because an EBS volume can only be attached to a single EC2 instance at a time (except for multi-attach EBS, which is limited to specific instance types and is not designed for shared file storage across AZs); attaching the same EBS volume to multiple instances is not supported.

783
Multi-Selectmedium

A company runs a media transcoding service on Amazon EC2 instances behind an Application Load Balancer. The workload is steady at 60% CPU utilization from 08:00 to 18:00 local time on weekdays and drops to under 5% overnight and on weekends. A solutions architect must reduce compute costs without changing the application code or degrading transcoding throughput during peak hours. (Choose two.)

Select 2 answers
A.Replace all On-Demand instances with Spot Instances in a single Availability Zone.
B.Move the transcoding workload to a larger instance type so fewer instances are needed at peak.
C.Enable detailed CloudWatch monitoring at one-minute resolution for every instance in the group.
D.Configure a scheduled scaling action on the Auto Scaling group to reduce desired capacity outside the 08:00-18:00 weekday window.
E.Purchase a 1-year Compute Savings Plan covering the baseline instance usage that runs every day.
AnswersD, E

Because demand is highly predictable by time of day, a scheduled scaling action can lower the minimum and desired capacity during nights and weekends so the group stops paying for idle instances. This matches capacity to the known traffic pattern and works alongside a commitment-based discount that covers only the always-on baseline.

Why this answer

The workload has two distinct cost components: an always-on baseline and a predictable daily peak. A Compute Savings Plan discounts the continuous baseline usage across any instance family or Region, while scheduled scaling trims the fleet during the known low-traffic windows. Together they reduce spend without touching application code, preserving peak throughput, and avoiding the interruption risk that Spot would introduce.

Exam trap

The trap here is assuming that any commitment-based discount must cover the entire fleet, when in fact the optimal design commits only to the always-on baseline and lets scheduled scaling handle the variable portion.

784
MCQeasy

An internal web application must require encrypted client connections. The company currently has an ALB listener on port 80 (HTTP), and users can access the application over plain HTTP. What is the best change to ensure all client traffic uses HTTPS?

A.Configure an HTTPS (port 443) listener using an ACM certificate and update the port 80 listener to redirect to HTTPS (or to block plain HTTP requests).
B.Enable S3 default encryption so HTTP requests are automatically encrypted in transit.
C.Set the application to encrypt data only after it is received by the ALB.
D.Rely on WAF alone to encrypt HTTP traffic.
AnswerA

Client-to-ALB encryption is enforced by terminating TLS on an ALB HTTPS listener. Redirecting or blocking HTTP on port 80 ensures clients cannot successfully establish plaintext HTTP sessions, so all viable paths use HTTPS end-to-end between the client and the load balancer.

Why this answer

It uses an HTTPS listener on port 443 with an ACM certificate to enforce encrypted client connections, and redirecting HTTP (port 80) traffic to HTTPS ensures all traffic is encrypted in transit. This is the standard AWS best practice for enforcing HTTPS on an ALB, as it directly controls the listener behavior at the load balancer level without requiring application changes.

Exam trap

The trap here is that candidates may confuse encryption at rest (S3 default encryption) with encryption in transit, or assume that WAF or post-receipt encryption can secure the initial client connection, when only a properly configured HTTPS listener with a redirect from HTTP can enforce encrypted client connections.

How to eliminate wrong answers

Option B is wrong because S3 default encryption applies to data at rest in S3 buckets, not to data in transit over HTTP; it cannot encrypt client-to-ALB traffic. Option C is wrong because encrypting data after it is received by the ALB means the initial client-to-ALB leg remains in plaintext HTTP, failing the requirement for encrypted client connections. Option D is wrong because AWS WAF is a web application firewall that inspects HTTP/HTTPS traffic but does not perform encryption; it cannot encrypt plain HTTP traffic.

785
MCQmedium

A microservice runs in private subnets with no NAT gateway. It must retrieve a secret from AWS Secrets Manager. Security requires that traffic to Secrets Manager stays within AWS’s private network (no public internet egress). The IAM role already grants secretsmanager:GetSecretValue for the needed secret. What is the best network setup to meet the requirement?

A.Create an Interface VPC Endpoint for Secrets Manager (com.amazonaws.<region>.secretsmanager) and allow it via the endpoint security group; optionally enable private DNS.
B.Create an S3 Gateway VPC endpoint and use it for Secrets Manager requests because both services use HTTPS.
C.Assign a public IP address to the tasks so they can call Secrets Manager over the internet without NAT.
D.Change the route table to send all 0.0.0.0/0 traffic directly to an Internet Gateway.
AnswerA

Interface VPC Endpoints provide private IP connectivity from the VPC to the Secrets Manager service without routing through a NAT gateway or an Internet Gateway. The calls remain within AWS networking and still use standard TLS to the service endpoint.

Why this answer

An Interface VPC Endpoint (AWS PrivateLink) for Secrets Manager allows the microservice to access the secret privately without traversing the public internet. Since the subnet has no NAT Gateway and no public IP, this is the only way to keep traffic within the AWS network. Enabling private DNS ensures the standard Secrets Manager endpoint resolves to the private IP of the endpoint, eliminating the need for route table changes.

Exam trap

The trap here is that candidates often confuse Gateway Endpoints (which only work for S3 and DynamoDB) with Interface Endpoints (which are needed for Secrets Manager and most other AWS services), leading them to incorrectly select option B.

How to eliminate wrong answers

Option B is wrong because S3 Gateway VPC endpoints are specific to Amazon S3 and cannot be used for Secrets Manager requests; Secrets Manager requires an Interface endpoint (powered by PrivateLink), not a Gateway endpoint. Option C is wrong because assigning a public IP address would route traffic over the public internet, violating the requirement that traffic stays within AWS’s private network. Option D is wrong because sending all 0.0.0.0/0 traffic to an Internet Gateway would force traffic out to the public internet, which is not allowed, and the subnet has no NAT Gateway to enable return traffic.

786
MCQeasy

A new feature stores user events in DynamoDB. Each event must be fetched by user_id and sorted by event_time. The team expects many different users and wants to avoid a single hot partition. Which partition key design is best?

A.Use a constant partition key value (for example, partition_key='events') and store user_id as an attribute.
B.Use user_id as the partition key and event_time as the sort key.
C.Use event_time as the partition key and user_id as an attribute to query later.
D.Use a randomly generated UUID as the partition key and query by user_id using a full table scan.
AnswerB

Using user_id as the partition key spreads data across many partitions based on user distribution. event_time as the sort key supports efficient range queries and retrieving events in time order per user. This design matches the stated access pattern and reduces hot partition likelihood.

Why this answer

Using user_id as the partition key evenly distributes writes across partitions, avoiding hot spots, while event_time as the sort key enables efficient retrieval of events for a specific user in chronological order. DynamoDB's query operation can then fetch all events for a given user_id sorted by event_time without scanning.

Exam trap

The trap here is that candidates may choose a constant partition key (Option A) thinking it simplifies queries, not realizing it creates a single hot partition that defeats DynamoDB's scalability.

Why the other options are wrong

A

Using a constant partition key like 'events' would cause all data to land on a single partition, creating a hot partition and defeating the purpose of DynamoDB's distributed architecture.

C

Using event_time as the partition key would create a hot partition for a specific time range (e.g., all events at the same second), causing throttling and uneven load distribution, which fails to avoid hot partitions as required.

When would these options actually be correct?

A

If the question required storing a small, fixed dataset (e.g., configuration data) where all items need to be retrieved together and write throughput is negligible, a constant partition key would be acceptable.

C

A question where the access pattern requires retrieving all events within a specific time window (e.g., 'get all events between 10:00 and 10:05') and the workload has low write volume, so a time-based partition key is acceptable.

Why candidates pick the wrong answer

A

Candidates may think a constant partition key simplifies queries (e.g., scanning all events) without realizing it creates a bottleneck for high-traffic workloads.

C

Candidates may think that sorting by event_time is best achieved by making it the partition key, overlooking that partition key design must ensure even distribution, not just sorting capability.

787
Multi-Selecthard

A log archive has old unattached EBS volumes and many stale snapshots. Which two actions reduce storage cost without affecting running instances? The architecture review board prefers a managed AWS-native control.

Select 2 answers
A.Stop all EC2 instances in the account
B.Disable CloudTrail logging
C.Delete unattached EBS volumes after verifying they are no longer needed
D.Apply snapshot lifecycle policies to expire obsolete snapshots
AnswersC, D

Unattached EBS volumes are billed as block storage regardless of whether they are mounted to a running instance. After confirming the volume is not needed for future use or as a boot source, deleting it releases the allocated storage and immediately stops accruing charges. For safety, you can first take a final snapshot, then delete the volume to preserve data recovery options without ongoing volume costs—this directly eliminates the recurring per-GiB-month expense for orphaned volumes.

Why this answer

Deleting unattached EBS volumes eliminates storage costs for volumes that are not in use, and since they are not attached to any running instance, this action does not affect running instances. Option D is correct because applying snapshot lifecycle policies (e.g., using Amazon Data Lifecycle Manager) automates the expiration of obsolete snapshots, reducing storage costs without impacting running instances. Both actions are managed AWS-native controls, aligning with the architecture review board's preference.

Exam trap

The trap here is that candidates may confuse stopping instances (which does not delete volumes) with deleting unattached volumes, or they may think disabling CloudTrail reduces storage costs, but CloudTrail logs are stored in S3 and are unrelated to EBS volume or snapshot storage charges.

788
Multi-Selectmedium

A company is designing a high-performance web application that serves static and dynamic content to a global user base. The application runs on Amazon EC2 instances behind an Application Load Balancer (ALB). The static assets are stored in an S3 bucket. Which three architecture decisions will improve performance and reduce latency for users? (Choose three.)

Select 3 answers
.Place the EC2 instances in a single Availability Zone to reduce network latency.
.Use Amazon CloudFront to cache both static and dynamic content at edge locations.
.Integrate the ALB with AWS Global Accelerator to route traffic over the AWS global network.
.Use a larger EC2 instance type with higher network bandwidth, such as the c5n or m5n family.
.Enable S3 Transfer Acceleration on the bucket for faster downloads.
.Use an Amazon RDS Multi-AZ database for read replicas to offload read traffic.

Why this answer

Amazon CloudFront caches both static and dynamic content at edge locations, reducing latency by serving content from locations closer to users. AWS Global Accelerator improves performance by routing traffic over the AWS global network instead of the public internet, reducing jitter and latency. Larger EC2 instance types like c5n or m5n provide higher network bandwidth, which reduces network bottlenecks for high-traffic applications.

Exam trap

The trap here is that candidates may confuse S3 Transfer Acceleration as a solution for faster downloads, when it only accelerates uploads, or think Multi-AZ RDS provides read scaling, when it is for failover only.

789
MCQmedium

A production log archive runs continuously on EC2 with predictable usage for the next three years. The team wants a discount while retaining some instance-family flexibility. What should they buy?

A.S3 Intelligent-Tiering
B.Dedicated Instances
C.Compute Savings Plan
D.Spot Instances only
AnswerC

A Compute Savings Plan commits to a fixed hourly spend for one or three years, applying discounted rates across EC2 instance families, sizes, and Regions, delivering the required discount while preserving the instance-family flexibility the team needs.

Why this answer

The Compute Savings Plan (C) is correct because it offers a discount (up to 66%) in exchange for a commitment to a consistent amount of compute usage (measured in $/hour) for a 1- or 3-year term, while allowing flexibility to change instance families, sizes, OS, tenancy, and even regions within EC2, Fargate, and Lambda. This matches the requirement of predictable usage for three years with instance-family flexibility, unlike Reserved Instances which lock to a specific instance family.

Exam trap

The trap here is that candidates often confuse Compute Savings Plans with Reserved Instances, assuming that any long-term discount requires locking into a specific instance family, but Compute Savings Plans provide both the discount and the flexibility to change instance families, which is the key differentiator tested in this question.

How to eliminate wrong answers

Option A is wrong because S3 Intelligent-Tiering is a storage class for objects in Amazon S3 that optimizes costs by moving data between access tiers based on changing access patterns; it has nothing to do with EC2 compute discounts or instance-family flexibility. Option B is wrong because Dedicated Instances are EC2 instances that run on hardware dedicated to a single customer, providing physical isolation but no discount or flexibility benefit; they are a billing/tenancy option, not a discount program. Option D is wrong because Spot Instances only offer significant discounts but are interruptible with a 2-minute termination notice, making them unsuitable for a production log archive that must run continuously for three years without interruption.

790
MCQeasy

Based on the exhibit, which AWS feature should the team use to minimize network latency between EC2 instances that exchange messages very frequently?

A.Use a spread placement group to maximize instance separation across hardware.
B.Use a cluster placement group to place instances close together.
C.Use a partition placement group to distribute instances across many partitions.
D.Use multiple Auto Scaling groups to spread traffic across more subnets.
AnswerB

A cluster placement group is designed for workloads that need very low network latency and high packet-per-second performance between instances. The exhibit describes frequent small-message traffic and a need for the lowest possible latency, which makes a cluster placement group the right choice. It keeps instances physically close in the AWS network for faster communication.

Why this answer

A cluster placement group is the correct choice because it places EC2 instances in a low-latency, high-bandwidth network within a single Availability Zone. This minimizes network latency between instances that exchange messages very frequently, as the instances are physically close together and can communicate using up to 10 Gbps of network throughput for most instance types.

Exam trap

The trap here is that candidates often confuse placement group types, assuming a spread placement group is for performance when it is actually designed for high availability and fault isolation, not low latency.

Why the other options are wrong

A

A spread placement group maximizes physical separation to reduce correlated failures, but it does not minimize network latency; in fact, it increases latency by placing instances farther apart.

C

A partition placement group distributes instances across logical partitions to isolate failures, but it does not minimize network latency; it may even increase latency due to physical separation.

D

Using multiple Auto Scaling groups to spread traffic across more subnets does not minimize network latency between EC2 instances; it may even increase latency by distributing instances across different availability zones or subnets, whereas the question requires placing instances close together for low-latency communication.

When would these options actually be correct?

A

When the question asks for a placement group that reduces the risk of simultaneous failures from shared hardware (e.g., for a small number of critical instances that must be isolated from each other).

C

When the question asks for fault isolation for large distributed workloads (e.g., Hadoop, Kafka) where you need to ensure that failures in one partition do not affect others, and low latency is not the primary concern.

D

This option would be correct in a scenario where the goal is to increase fault tolerance and high availability by distributing instances across multiple availability zones, and the question specifically asks for a solution to handle traffic spikes or ensure resilience, not minimize latency.

Why candidates pick the wrong answer

A

Candidates may confuse 'spread' with 'low latency' because spreading instances could reduce contention, but the primary goal of spread placement is fault isolation, not network performance.

C

Candidates may confuse 'partition' with 'cluster' or think that distributing instances reduces latency by avoiding congestion, but partition groups are designed for resilience, not low latency.

D

Candidates might think that spreading traffic across more subnets reduces network congestion and thus latency, but they overlook that the primary factor for low latency between frequently communicating instances is physical proximity, not traffic distribution.

791
MCQeasy

A team uses an S3 bucket to store important customer-generated exports. They need protection against accidental overwrites and also want copies of the data in another AWS Region for disaster recovery. Which S3 configuration best satisfies both requirements?

A.Enable S3 lifecycle policies to automatically move objects to Glacier after 30 days only.
B.Enable S3 versioning and configure Cross-Region Replication to a destination bucket in another Region.
C.Disable all versioning and rely on AWS Backup to restore objects from a scheduled backup window.
D.Enable S3 Block Public Access and SSE-S3 encryption, without using versioning or replication.
AnswerB

Enabling S3 versioning preserves every version of an object, so accidental overwrites or deletes can be undone by restoring a prior version or removing a delete marker. Cross-Region Replication then asynchronously copies new and updated objects to a bucket in another Region, providing a geographically separate copy for disaster recovery. Together these features directly address both object-level corruption and Region-level failures, making them the correct solution.

Why this answer

Enabling S3 versioning protects against accidental overwrites by preserving all object versions, allowing recovery of previous versions. Configuring Cross-Region Replication (CRR) automatically replicates objects to a destination bucket in another AWS Region, providing disaster recovery by maintaining a copy of the data in a separate geographic location.

Exam trap

The trap here is that candidates may think lifecycle policies or AWS Backup alone can handle both accidental overwrites and disaster recovery, but they fail to address the real-time protection and cross-region copy requirements that versioning and CRR specifically provide.

Why the other options are wrong

A

Lifecycle policies to Glacier only address storage cost optimization, not protection against accidental overwrites or cross-region disaster recovery.

C

AWS Backup does not prevent accidental overwrites; it only provides scheduled backups. Without versioning, overwritten objects are permanently lost until the next backup, and recovery point objectives may not align with real-time protection.

D

Block Public Access and SSE-S3 encryption protect against unauthorized access and encrypt data at rest, but they do not prevent accidental overwrites or provide cross-region disaster recovery copies.

When would these options actually be correct?

A

A question asking for cost-effective long-term archival of infrequently accessed data, with no requirement for versioning or cross-region replication, would make this correct.

C

If the question required a centralized backup solution across multiple AWS services (e.g., EC2, RDS, and S3) with a defined retention policy and compliance auditing, AWS Backup would be the correct answer.

D

An exam question asking for the best way to secure an S3 bucket from public access and ensure data encryption at rest, without mentioning versioning or replication requirements, would make this option correct.

Why candidates pick the wrong answer

A

Candidates may confuse lifecycle management with data protection, assuming moving to Glacier provides backup or recovery capabilities.

C

Candidates may assume that AWS Backup offers the same protection as versioning for overwrites, or they overestimate the frequency of backups, not realizing that versioning provides immediate recovery without relying on backup schedules.

D

Candidates may mistakenly think that security measures like Block Public Access and encryption are sufficient for data protection and disaster recovery, overlooking the need for versioning and replication.

792
MCQmedium

A media company runs a video-transcoding fleet on Amazon EC2 instances that read source files from an Amazon S3 bucket and write output to a second bucket. The fleet is spread across three Availability Zones in one Region, and instances are launched by an Auto Scaling group. The company needs the architecture to survive the loss of an entire Availability Zone without losing in-flight transcoding work or requiring manual intervention. Which combination of design elements should a solutions architect implement to meet these requirements?

A.Deploy the Auto Scaling group across three Availability Zones, make transcoding jobs idempotent and store progress in Amazon DynamoDB, and have instances poll an Amazon SQS queue for work so that unfinished jobs are retried by healthy instances.
B.Deploy the Auto Scaling group in a single Availability Zone with a spot fleet, use an Amazon EBS volume attached to each instance to persist transcoding progress, and enable EBS snapshots every five minutes to another zone.
C.Deploy the Auto Scaling group across three Availability Zones, place a Network Load Balancer in front of the instances, and configure the load balancer to retry failed transcoding requests against the same instance until the zone recovers.
D.Deploy the Auto Scaling group across three Availability Zones, store transcoding state in an Amazon S3 bucket configured with S3 Cross-Region Replication, and rely on the S3 Standard storage class for automatic recovery.
AnswerA

Spreading the Auto Scaling group across three Availability Zones means instances in surviving zones continue running when one zone fails. Because work is pulled from an SQS queue and progress is tracked in DynamoDB, a job interrupted in the failed zone becomes visible again after its visibility timeout and is retried by a healthy instance, so no manual intervention is required.

Why this answer

Resilience across an Availability Zone failure requires both compute capacity in the surviving zones and durable, externalized job state. Distributing the Auto Scaling group across three zones keeps instances running, while an SQS queue with visibility timeouts and DynamoDB progress tracking lets interrupted jobs be reclaimed and retried automatically. This removes any dependency on the failed zone's instances or storage.

Exam trap

The trap here is assuming that storing source and output objects in Amazon S3 is sufficient for workload resilience, when S3 durability says nothing about resuming in-flight compute work after a zone failure.

793
Multi-Selectmedium

A solutions architect is designing a high-performance architecture for a read-heavy web application backed by Amazon RDS for MySQL. The database is currently a single db.r6g.4xlarge instance that is CPU-bound during peak hours, and the application performs many repeated identical read queries. The architect must improve read scalability and reduce load on the primary instance. (Choose two.)

Select 2 answers
A.Enable RDS Performance Insights and set up a longer retention period for the performance history.
B.Add an Amazon ElastiCache for Redis cluster in front of the database and cache the results of repeated queries.
C.Convert the instance to a larger db.r6g.16xlarge size to increase memory and CPU.
D.Create one or more read replicas and direct read-only queries to the replica endpoints.
E.Enable Multi-AZ with a standby instance and point read queries at the standby endpoint.
AnswersB, D

Caching the results of frequently repeated identical queries in ElastiCache for Redis means the application can serve them from memory without touching RDS at all, which removes a large share of read traffic. This complements read replicas and is especially effective for hot keys that would otherwise hit the primary on every request.

Why this answer

Read replicas add horizontal read capacity by serving queries from separate instances, while an ElastiCache for Redis cluster absorbs repeated identical reads in memory before they reach the database. Together they reduce the CPU load on the primary from both directions: fewer queries arrive, and the ones that do can be served by replicas instead of the writer.

Exam trap

The trap here is believing that a Multi-AZ standby can serve read traffic, when it exists solely for high availability failover and never handles application queries.

794
Multi-Selectmedium

A company runs an internal API on Amazon EC2 instances in a private subnet. Clients in an on-premises data center must reach the API over a private connection, and the security team wants to inspect and filter the traffic using AWS managed security appliances before it reaches the application. The company has already established an AWS Site-to-Site VPN to a transit gateway. Which two actions should the security engineer take to route and inspect the traffic? (Choose two.)

Select 2 answers
A.Create a VPC attachment on the transit gateway for the VPC that hosts the API and the inspection appliances.
B.Configure the VPC route tables so that traffic from the VPN attachment is directed to the subnet hosting the inspection appliances, and traffic from the appliances is directed to the API subnet.
C.Attach an internet gateway to the VPC and advertise the VPC CIDR to on-premises over the VPN so that return traffic uses the public path.
D.Enable VPC flow logs on the API subnet and use Amazon GuardDuty findings to block malicious sources at the network layer.
E.Replace the Site-to-Site VPN with an AWS Direct Connect connection because a transit gateway cannot route traffic from a VPN attachment into a VPC.
AnswersA, B

The transit gateway must have a VPC attachment before it can route traffic between the VPN and the VPC subnets. Without that attachment the VPN tunnel terminates at the transit gateway but has no path into the VPC, so packets from on-premises never reach the inspection appliances or the API instances.

Why this answer

Traffic from the VPN reaches the transit gateway, which needs a VPC attachment to hand packets into the VPC. Once inside, route tables must send the flows through the subnet that hosts the inspection appliances and then on to the API subnet, with a symmetric return path, so that every packet is evaluated by the security layer before delivery.

Exam trap

The trap here is assuming that attaching the VPN to the transit gateway is sufficient and that traffic will be inspected automatically, when inspection only occurs if VPC route tables explicitly steer packets through the appliance subnet.

795
MCQhard

A media archive needs low-latency full-text search across product descriptions and filtered attributes. Which managed service is most suitable?

A.AWS Config
B.Amazon OpenSearch Service
C.Amazon EFS
D.Amazon SQS
AnswerB

Amazon OpenSearch Service is a fully managed search and analytics engine based on the open-source OpenSearch project. It uses an inverted index to enable low-latency full-text search, supports complex query syntax, analyzers, tokenization, and relevance scoring, making it ideal for searching large amounts of text such as media metadata, transcripts, or subtitles. Its ability to ingest and index data from services like S3 or Kinesis and return results in near real time directly addresses the media archive's requirement.

Why this answer

Amazon OpenSearch Service is the correct choice because it provides a managed, scalable solution for full-text search and real-time analytics on large volumes of data. It supports low-latency queries across product descriptions and filtered attributes through its inverted index and query DSL, making it ideal for media archive search use cases.

Exam trap

The trap here is that candidates may confuse AWS Config's resource tracking or SQS's message handling with search capabilities, but neither provides the indexing and query engine required for full-text search.

How to eliminate wrong answers

Option A is wrong because AWS Config is a service for auditing and evaluating resource configurations against desired policies, not for full-text search or indexing of data. Option C is wrong because Amazon EFS is a scalable file storage service for Linux-based workloads, lacking any built-in search or indexing capabilities for text content. Option D is wrong because Amazon SQS is a fully managed message queuing service for decoupling application components, not designed for storing or searching data.

796
MCQmedium

A media company hosts a public-facing web application on Amazon EC2 instances behind an Application Load Balancer. The security team wants to protect the application from common web exploits such as SQL injection and cross-site scripting, and also wants to rate-limit requests from specific IP addresses that exhibit abusive behavior. Which combination of AWS services should a solutions architect recommend?

A.AWS WAF attached to the Application Load Balancer with managed rule groups and rate-based rules.
B.AWS Shield Advanced with automatic application layer DDoS mitigation enabled on the ALB.
C.Amazon GuardDuty with suppression rules configured for the EC2 instances and ALB.
D.AWS Network Firewall deployed in a subnet in front of the ALB with stateful rules for HTTP inspection.
AnswerA

AWS WAF integrates directly with Application Load Balancers and supports managed rule groups that block common exploits like SQL injection and XSS. Rate-based rules can automatically block or challenge IP addresses that exceed a defined request threshold within a five-minute window. This single service satisfies both the exploit protection and the rate-limiting requirement with minimal operational overhead and native ALB integration.

Why this answer

AWS WAF is the managed application-layer firewall that attaches to Application Load Balancers, Amazon CloudFront, and API Gateway. Managed rule groups provide prebuilt protections against SQL injection, XSS, and other OWASP threats, while rate-based rules track request counts per IP and block offenders automatically. This directly addresses both the exploit protection and the abusive IP rate-limiting requirements with a single integrated service.

Exam trap

The trap here is confusing DDoS protection or network-layer firewalls with application-layer exploit filtering, which only AWS WAF provides for ALBs.

797
MCQmedium

Based on the exhibit, what is the best way to let private EC2 instances reach Amazon S3 and AWS Systems Manager without sending traffic through the internet or a NAT gateway?

A.Create a gateway endpoint for S3 and interface endpoints for Systems Manager, EC2Messages, and SSMMessages.
B.Add a more permissive security group rule allowing outbound 0.0.0.0/0 on all ports.
C.Replace the NAT gateway with a network ACL that allows ephemeral ports to the internet.
D.Move the instances to public subnets so they can reach AWS services directly.
AnswerA

This keeps traffic on the AWS network and avoids NAT or internet traversal. S3 uses a gateway endpoint, while Systems Manager needs interface endpoints for the control and messaging services that Session Manager depends on. It directly addresses both the S3 download problem and the missing Session Manager connectivity in a private subnet design.

Why this answer

Gateway endpoints for S3 allow private EC2 instances to access S3 via AWS's private network without traversing the internet or a NAT gateway, using prefix lists and route table entries. Interface endpoints for Systems Manager, EC2Messages, and SSMMessages provide private connectivity to AWS Systems Manager via PrivateLink, enabling secure instance management without public IPs or NAT.

Exam trap

The trap here is that candidates often assume all AWS services can be accessed via a single endpoint type, but S3 requires a gateway endpoint (route table-based) while Systems Manager and its sub-services require interface endpoints (PrivateLink-based), and failing to create all three (including EC2Messages and SSMMessages) will break Systems Manager functionality.

Why the other options are wrong

B

This option does not address the requirement to avoid sending traffic through the internet or a NAT gateway; a permissive security group rule still routes traffic via the internet or NAT gateway, not through private VPC endpoints.

C

Network ACLs cannot replace NAT gateways for outbound internet access; they are stateless and only filter traffic, not provide NAT. Private instances still need a NAT device to reach the internet, and Systems Manager requires VPC endpoints, not just internet access.

D

Moving instances to public subnets exposes them directly to the internet, which violates the requirement to avoid sending traffic through the internet. Private instances should remain in private subnets and use VPC endpoints for secure, private connectivity to AWS services.

When would these options actually be correct?

B

This option would be correct in a scenario where the question asks for the simplest way to allow outbound internet access from private EC2 instances, without any restriction on using the internet or NAT gateway, and the goal is just to enable general outbound connectivity.

C

If the question asked how to restrict outbound traffic from a public subnet to the internet using only network ACLs (e.g., allow ephemeral ports while blocking other traffic), then replacing a NAT gateway with a properly configured network ACL could be correct, assuming instances have public IPs or are behind an internet gateway.

D

If the question asked for the simplest way to allow instances to reach the internet (not just AWS services) without a NAT gateway, and security concerns were not a factor, moving instances to public subnets with public IPs and appropriate security groups would enable direct internet access.

Why candidates pick the wrong answer

B

Candidates may think that allowing all outbound traffic via a security group is a quick fix, overlooking the specific requirement to avoid internet or NAT gateway paths and the need for private connectivity via endpoints.

C

Candidates may confuse network ACLs with NAT gateways, thinking ACLs can provide outbound connectivity by allowing ephemeral ports, or they may underestimate the need for NAT to translate private IPs.

D

Candidates may think that public subnets provide direct access to all AWS services, overlooking that public subnets still route traffic through the internet, and that private connectivity via VPC endpoints is more secure and cost-effective for services like S3 and Systems Manager.

798
MCQeasy

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The application stores user-uploaded images in an Amazon S3 bucket. Users report slow image upload times, especially from mobile devices in remote locations. The solutions architect needs to improve upload performance for these users. Which action should the architect take?

A.Configure the S3 bucket for cross-region replication to a bucket in a region closer to the users.
B.Enable S3 Transfer Acceleration on the bucket and configure the application to use the accelerated endpoint for uploads.
C.Use Amazon CloudFront with a custom origin pointing to the S3 bucket to cache uploads.
D.Increase the size of the EC2 instances to improve their network bandwidth for handling uploads.
AnswerB

S3 Transfer Acceleration uses Amazon CloudFront edge locations to accelerate uploads over long distances. When enabled, uploads go to the nearest edge location and are then routed to S3 over the AWS global network, reducing latency and improving throughput. This is ideal for mobile users in remote locations. The application must use the accelerated endpoint, which is a simple configuration change, to benefit from the acceleration.

Why this answer

S3 Transfer Acceleration leverages CloudFront edge locations to accelerate uploads to S3. Data is uploaded to the nearest edge location and then transferred to S3 over the optimized AWS network. This reduces latency and improves throughput for users far from the bucket's Region, making it the right choice for slow image uploads from remote mobile devices.

Exam trap

The trap here is thinking that a CDN like CloudFront accelerates uploads by default, when it is primarily for download acceleration; S3 Transfer Acceleration is the service specifically designed for faster uploads.

799
MCQmedium

A media company runs a fleet of EC2 instances using Auto Scaling across multiple instance families (for example, m-series and c-series) in a single region. The business wants to commit to steady usage for one year to reduce cost, but the application team must retain flexibility to switch instance families and scale up/down as demand changes. They need the cost-reduction approach that best matches this flexibility. Which option is the best fit?

A.Purchase Standard Reserved Instances tied to a specific instance family and region, so the application can only run on the selected family.
B.Purchase Compute Savings Plans so the commitment applies regardless of instance family changes within the selected scope.
C.Purchase Spot Instances for all capacity and disable On-Demand fallback to guarantee the lowest cost.
D.Rely only on On-Demand and reduce cost by using a CloudFront-only approach for all dynamic content.
AnswerB

Compute Savings Plans provide discounted pricing in exchange for a 1-year or 3-year commitment, while allowing flexibility across instance families/attributes within the scope (for example, region/account and covered usage). This aligns with Auto Scaling that may shift between instance families while maintaining steady overall compute usage.

Why this answer

Compute Savings Plans provide the most flexibility because they apply to any EC2 instance family (including m-series and c-series) within a region, automatically adjusting to instance family changes and scaling. This matches the requirement to commit to steady usage for one year while retaining the ability to switch families and scale up/down, offering up to 66% savings over On-Demand without locking the application to a specific instance type.

Exam trap

The trap here is that candidates often confuse Reserved Instances (which lock to a specific family) with Savings Plans (which offer family flexibility), leading them to choose Option A despite the requirement for instance family switching.

How to eliminate wrong answers

Option A is wrong because Standard Reserved Instances are tied to a specific instance family (e.g., m5.large) and region, which would prevent the application from switching to a different instance family (e.g., c-series) without incurring additional On-Demand costs or modification fees. Option C is wrong because Spot Instances can be interrupted with a 2-minute warning, making them unsuitable as the sole capacity source for a production workload that requires reliability; disabling On-Demand fallback would risk application downtime during Spot reclaimations. Option D is wrong because CloudFront is a content delivery network that caches static and dynamic content at edge locations, but it does not reduce the cost of running EC2 instances for compute workloads; relying solely on On-Demand without a commitment discount would not achieve the desired cost reduction.

800
MCQhard

Based on the exhibit, a serverless checkout API is implemented in AWS Lambda and deployed in one Region. The function has a cold-start time of 700-900 ms on the first request after idle periods. Marketing launches a predictable traffic spike every weekday at 09:00 UTC, and the p95 latency target is under 150 ms during the first five minutes of the spike. What should the solutions architect do to meet the latency target while controlling cost?

A.Increase the Lambda memory size and leave concurrency at the default value.
B.Configure provisioned concurrency and scale it up before the predictable spike begins.
C.Put the Lambda function behind an Application Load Balancer so the load balancer absorbs the initialization delay.
D.Set reserved concurrency to the expected peak so Lambda will pre-create execution environments.
AnswerB

Provisioned concurrency keeps a specified number of execution environments fully initialized and idle, ready to serve invocations in milliseconds instead of incurring a cold start. Because the spike is predictable, you can configure scheduled scaling to raise provisioned capacity to the expected peak before 09:00 UTC and lower it afterward, so users see low latency without paying idle costs all day. This is the only approach that directly eliminates the initialization delay by pre-warming the sandboxes.

Why this answer

Provisioned concurrency pre-warms a specified number of execution environments so that the Lambda function has zero cold-start latency when invoked. By scheduling the provisioned concurrency to scale up before the 09:00 UTC spike, the function can serve the first requests within the 150 ms p95 latency target, while the scheduled scaling down after the spike controls cost by releasing unused capacity.

Exam trap

AWS often tests the distinction between provisioned concurrency (which pre-warms environments to eliminate cold starts) and reserved concurrency (which only caps the maximum concurrent executions without affecting cold-start behavior).

How to eliminate wrong answers

Option A is wrong because increasing memory size can reduce cold-start time but cannot eliminate it entirely, and the cold-start of 700-900 ms far exceeds the 150 ms target; default concurrency does not pre-warm environments. Option C is wrong because an Application Load Balancer does not absorb initialization delay—it only distributes requests to the Lambda function, which still experiences cold starts. Option D is wrong because reserved concurrency limits the maximum number of concurrent executions but does not pre-create execution environments; it prevents scaling beyond a limit but does not reduce cold-start latency.

801
Multi-Selectmedium

A solutions architect is designing a high-performance architecture for a web application that serves static content from Amazon S3 and dynamic content from an Application Load Balancer. The application must deliver low latency to users across multiple continents and reduce origin load. The team wants to use Amazon CloudFront. Which two actions should the architect take to meet these requirements? (Choose two.)

Select 2 answers
A.Use Lambda@Edge to rewrite all requests to a single origin and compress responses, eliminating the need for multiple cache behaviors and origins.
B.Configure the S3 bucket as a CloudFront origin with an Origin Access Control (OAC) and set the bucket policy to allow only the CloudFront distribution to access the objects.
C.Create a second CloudFront distribution with a different price class and use Route 53 latency-based routing to direct users to the distribution with the lowest latency.
D.Enable CloudFront caching for static assets with a long TTL and configure the Application Load Balancer as a second origin with cache behaviors that forward necessary headers for dynamic requests.
E.Set the Application Load Balancer as a custom origin, enable caching for dynamic content with a long default TTL, and forward all headers and cookies to the origin.
AnswersB, D

Using OAC lets CloudFront securely access the S3 bucket without making objects public, and the bucket policy restricts direct access. This improves security and allows CloudFront to cache static content at edge locations, reducing latency for global users and lowering the load on the S3 origin because repeated requests are served from cache.

Why this answer

A CloudFront distribution with multiple origins and cache behaviors lets static S3 content be cached at edge locations with long TTLs while dynamic ALB traffic forwards only necessary headers. Using OAC secures S3 access. Together, these reduce latency for global users and lower origin load.

Multiple distributions or request rewriting do not provide the same performance and security benefits.

Exam trap

The trap here is thinking that multiple CloudFront distributions with latency-based routing are needed for global performance, when CloudFront's anycast network already routes users to the nearest edge location automatically.

802
MCQhard

A financial analytics team runs a read-heavy workload on Amazon Aurora MySQL. The primary instance is experiencing high CPU during end-of-day reporting, and read replicas are lagging by several seconds. The application requires strong read consistency for account balances but can tolerate eventual consistency for historical reports. Which change should a solutions architect make to improve performance while meeting consistency requirements?

A.Enable Aurora Auto Scaling for read replicas and direct all queries, including balance lookups, to the reader endpoint to distribute load evenly.
B.Create a second Aurora cluster that replicates from the primary using binary log replication, and direct all reporting queries to the second cluster's writer endpoint.
C.Route balance queries to the primary instance endpoint and historical report queries to the reader endpoint, and enable Aurora Auto Scaling to add read replicas based on replica utilization.
D.Increase the primary instance size and enable the query cache, then send all read traffic to the primary to ensure strong consistency for every query.
AnswerC

Balance lookups need strong consistency, so they must go to the primary. Historical reports tolerate eventual consistency and can be sent to the reader endpoint. Aurora Auto Scaling adds replicas when readers are busy, increasing aggregate read throughput and reducing lag. This separation preserves correctness while offloading the primary from reporting traffic.

Why this answer

Strongly consistent reads must target the primary instance, while eventually consistent reporting can be served by Aurora read replicas. Aurora Auto Scaling dynamically adds replicas based on load, increasing read capacity and reducing lag. This split preserves correctness for account balances and offloads reporting from the primary, improving overall performance without compromising consistency.

Exam trap

The trap here is believing that the reader endpoint can always serve balance lookups as long as replicas are added, when in fact replica lag means those reads may be stale and violate a strong consistency requirement.

803
Multi-Selectmedium

A company is designing a disaster recovery plan for a critical application hosted on AWS. The application runs on EC2 instances with data stored in Amazon EBS volumes and Amazon S3. The recovery time objective (RTO) is 15 minutes, and the recovery point objective (RPO) is 1 hour. Which three strategies would help meet these objectives? (Choose three.)

Select 3 answers
.Use AWS Backup to create hourly snapshots of EBS volumes and copy them to a different AWS Region.
.Pre-provision EC2 instances in the disaster recovery region and keep them running 24/7.
.Replicate critical data to S3 in the disaster recovery region using S3 Cross-Region Replication (CRR).
.Store Amazon Machine Images (AMIs) in the source region and use AWS Lambda to copy them after a disaster.
.Configure Amazon Route 53 with a failover routing policy and health checks to redirect traffic to the DR region.
.Set up an AWS Direct Connect link between the primary and DR regions for faster data transfer.

Why this answer

AWS Backup can create hourly snapshots of EBS volumes and copy them to a different AWS Region, meeting the 1-hour RPO by ensuring backups are taken every hour. S3 Cross-Region Replication (CRR) asynchronously replicates objects to a bucket in another region, keeping data synchronized within minutes and supporting the RPO. Amazon Route 53 with a failover routing policy and health checks can automatically redirect traffic to the DR region within seconds to minutes, enabling the 15-minute RTO by quickly failing over to pre-prepared infrastructure.

Exam trap

The trap here is that candidates may confuse operational readiness (like pre-provisioning instances) with a specific strategy that directly contributes to meeting RTO/RPO, or they may think Direct Connect is a disaster recovery strategy when it is merely a connectivity option that does not automate failover or data replication.

804
MCQmedium

A company needs to implement session management for a web application. Sessions must persist across multiple EC2 instances, survive EC2 failures, and be accessible with sub-millisecond latency. Sessions must also be sortable by last-access time to expire the oldest sessions first. Which caching solution should a solutions architect recommend?

A.Amazon ElastiCache for Memcached with session data stored as key-value pairs
B.Amazon DynamoDB with TTL enabled for session expiration
C.Amazon ElastiCache for Redis with sessions stored as sorted sets
D.ElastiCache for Redis with sticky sessions enabled on the Application Load Balancer
AnswerC

ElastiCache for Redis is the correct solution because it stores sessions in memory for sub-millisecond latency and uses sorted sets to maintain an ordered score per session, ideal for tracking last-access time. Multiple replicas with Multi-AZ deployment provide automatic failover, and because the session store is external to EC2 instances, every instance can serve any session without state affinity. ZADD/ZINCRBY and ZRANGE operations deliver exactly the required sorted-set semantics for session management.

Why this answer

Amazon ElastiCache for Redis satisfies all requirements: multi-instance session sharing (sessions stored externally), sub-millisecond latency, survival of EC2 failures (stored outside instances), and sorted sets (ZSET data structure) for ordering sessions by last-access score.

Memcached supports only simple key-value pairs — it cannot perform sorted set operations to order sessions by last-access time. Memcached also lacks replication, meaning a node failure loses all cached sessions.

Exam trap

Memcached and Redis are both ElastiCache engines, but they serve different needs. Any requirement involving sorted data, complex data structures, persistence, or replication eliminates Memcached. Redis sorted sets (ZSET) store members with numeric scores and support range queries — perfect for session expiry queues ordered by last-access timestamp.

Why the other options are wrong

A

Memcached supports only simple string key-value storage. It cannot perform sorted set operations to expire sessions by last-access time. Memcached also lacks replication — a node failure loses all cached sessions.

B

DynamoDB achieves single-digit millisecond latency, not sub-millisecond. DynamoDB also does not natively support sorted set operations without additional query complexity.

D

ALB sticky sessions pin a client to a specific EC2 instance. If that instance fails, the session is lost. Sticky sessions do not make session data redundant across instances — the opposite of what is required.

805
MCQhard

A security team stores sensitive documents in an Amazon S3 bucket that is encrypted with SSE-KMS using a customer managed key. An auditor requires that every object upload be traceable to the IAM principal that performed it and that the key's usage be independently auditable. The team also wants to prevent any principal, including account administrators, from reading objects without a corresponding key grant. Which configuration combination meets these requirements?

A.Enable AWS CloudTrail data events for the S3 bucket, enable KMS key rotation, and apply a bucket policy that requires the aws:SecureTransport condition.
B.Enable S3 server access logging to a separate bucket, enable S3 default encryption with SSE-S3, and attach a bucket policy that denies unencrypted uploads.
C.Enable AWS CloudTrail data events for the S3 bucket, enable CloudTrail management events for KMS, and ensure the key policy delegates usage only to explicitly named roles.
D.Enable S3 Object Lock in compliance mode, enable CloudTrail management events for S3, and use a key policy that grants kms:Decrypt to the account root user.
AnswerC

S3 data events record the identity of the caller for object-level operations such as PutObject and GetObject, while KMS management events capture Encrypt, Decrypt, and key policy changes. Restricting the key policy to named roles means an administrator without an explicit key grant cannot decrypt the data even if the bucket policy allows the S3 action, providing independent control over the key.

Why this answer

Object-level traceability requires CloudTrail data events on the bucket, and independent key auditing requires CloudTrail management events on KMS so that Encrypt and Decrypt calls appear in the log. Limiting the key policy to named roles creates a second authorization boundary, so possession of S3 permissions alone is insufficient to read the encrypted objects.

Exam trap

The trap here is believing that S3 server access logging or CloudTrail management events alone can attribute object uploads and prove key usage, when object-level activity requires data events and key usage requires KMS management events.

806
MCQmedium

A service consumes messages from an SQS queue. Recently, a new message format started failing validation in the consumer. The consumer catches the exception but cannot successfully process those messages without code changes. The team wants failed messages to be isolated for later investigation instead of being retried indefinitely. What should they configure?

A.Set the queue’s retention period to 1 minute and rely on messages expiring naturally.
B.Configure a dead-letter queue (DLQ) with a redrive policy and set maxReceiveCount so messages move after repeated failed receives.
C.Increase the visibility timeout to 7 days so failed messages cannot be retried.
D.Publish the same message again to SNS on every failure so a different subscriber might succeed.
AnswerB

A DLQ isolates “poison messages” that repeatedly fail processing. With a redrive policy, SQS tracks receives; once a message exceeds maxReceiveCount without successful processing, SQS moves it to the DLQ. This prevents infinite retries on the bad format while preserving the failed messages for debugging and code fixes.

Why this answer

A dead-letter queue (DLQ) with a redrive policy is the correct solution because it allows messages that repeatedly fail processing to be moved to a separate queue after exceeding the maxReceiveCount. This isolates problematic messages for later investigation without blocking the main queue or causing infinite retries. The consumer catches the exception, so the message is not deleted and is returned to the queue for redelivery; the DLQ ensures that after a configurable number of attempts, the message is redirected instead of being retried indefinitely.

Exam trap

The trap here is that candidates may think increasing the visibility timeout or relying on message expiration is sufficient, but they fail to understand that those approaches either affect all messages or only temporarily hide the message, whereas a DLQ provides a permanent, targeted isolation mechanism for repeatedly failing messages.

How to eliminate wrong answers

Option A is wrong because setting the retention period to 1 minute would cause all messages (including valid ones) to expire quickly, leading to data loss and not isolating only the failed messages. Option C is wrong because increasing the visibility timeout to 7 days would simply hide the message from consumers for that period, but after the timeout expires the message would become visible again and be retried, failing to isolate it permanently. Option D is wrong because publishing the same message to SNS on every failure would create an infinite loop of republishing, and SNS subscribers would also fail if they use the same validation logic, not solving the isolation requirement.

807
Multi-Selecthard

A solutions architect is reviewing a workload that runs on a fleet of Amazon EC2 instances in a single AWS Region. The application serves a global user base, and the team wants to reduce both data transfer costs and latency for users in Europe and Asia. The application is stateless and stores assets in Amazon S3. The team is also evaluating how to pay for the compute layer over the next three years, as usage is expected to be steady. Which two actions will reduce cost in this scenario? (Choose two.)

Select 2 answers
A.Enable S3 Transfer Acceleration on the bucket and direct all user downloads through the accelerated endpoint.
B.Purchase a three-year Compute Savings Plan with a full upfront payment for the EC2 compute.
C.Deploy an Amazon CloudFront distribution with the S3 bucket as the origin to cache assets at edge locations.
D.Move the S3 bucket to a Region closer to the European users and replicate back to the original Region.
E.Convert the EC2 instances to Spot Instances to eliminate compute charges entirely.
AnswersB, C

A Compute Savings Plan commits to a consistent amount of compute usage in exchange for a discount of up to 66 percent compared with On-Demand, and the longer the term and the larger the upfront payment, the greater the discount. Because usage is expected to be steady for three years, this commitment matches the demand and lowers the effective hourly cost of the fleet.

Why this answer

The workload has two distinct cost drivers: global asset delivery and steady compute. Caching assets at CloudFront edge locations reduces origin egress and internet transfer while improving latency for distant users. For the compute layer, a three-year Compute Savings Plan matches the expected steady usage and applies a significant discount over On-Demand.

Together these address both transfer and compute costs without sacrificing availability.

Exam trap

The trap here is treating S3 Transfer Acceleration as a cost-saving feature, when it actually adds a per-gigabyte charge on top of normal transfer and is intended for upload speed, not cheap global read delivery.

808
MCQeasy

A production Amazon RDS database has automated backups enabled with sufficient retention. At 10:30 UTC, a release corrupts specific rows. The issue is detected at 10:45 UTC. The team wants to restore the database state to before the corruption with minimal complexity. What should they do?

A.Perform a point-in-time restore (PITR) to a timestamp just before 10:30 UTC and create a restored DB instance/cluster.
B.Change the VPC route tables so the database restarts in a clean state.
C.Relaunch the same DB instance in the same Availability Zone and rely on caching to revert the changes.
D.Enable a DLQ on the database to store invalid SQL statements until the system is fixed.
AnswerA

PITR uses automated backups to restore the database to a specific point in time. Selecting a timestamp just before the corruption (for example, slightly before 10:30 UTC) restores the affected data state as it existed before the bad release.

Why this answer

Amazon RDS Point-in-Time Restore (PITR) allows you to restore the database to any second within the backup retention period, using automated backups and transaction logs. By restoring to a timestamp just before 10:30 UTC, you can recover the database to a state before the corruption occurred, creating a new DB instance/cluster with minimal complexity and no data loss from the uncorrupted period.

Exam trap

The trap here is that candidates may confuse database recovery methods with network or application-level fixes, or incorrectly assume that restarting or relaunching an instance will clear data changes, when in fact only a restore from backup or PITR can revert committed transactions.

How to eliminate wrong answers

Option B is wrong because changing VPC route tables affects network traffic routing, not database state or data integrity; it cannot revert corrupt rows or restart the database in a clean state. Option C is wrong because relaunching the same DB instance in the same Availability Zone does not revert data changes; it simply creates a new instance with the same underlying storage, which still contains the corrupt rows. Option D is wrong because a Dead Letter Queue (DLQ) is a concept for message queues (like Amazon SQS) to handle failed message processing, not a feature of Amazon RDS; it cannot store or revert SQL statements.

809
MCQmedium

A Multi-AZ Amazon RDS database experiences incorrect writes at 10:15 UTC due to a buggy release. The team detects the problem at 10:25 UTC. They want to restore the data to a known-good point around 10:15 UTC, and validate the recovered data, without taking the current production instance offline during the recovery process. What is the most appropriate AWS action?

A.Immediately reboot the RDS instance and rely on the reboot to roll back the bad writes.
B.Perform a point-in-time restore (PITR) to a new DB instance using a restore time around 10:15 UTC, then test the restored instance before cutting over.
C.Create a new Read Replica from the current primary and use it as the recovered database after applying reverse migrations.
D.Temporarily disable Multi-AZ to speed up storage rollback, then re-enable Multi-AZ.
AnswerB

PITR restores to a specific timestamp using backups and transaction logs. Importantly, it creates a recovered copy (typically a new DB instance), which allows validation and cutover decisions without stopping or directly impacting the existing production instance.

Why this answer

Amazon RDS point-in-time recovery (PITR) allows you to restore a DB instance to any second within the backup retention period, creating a new, independent DB instance. This lets you validate the recovered data without affecting the current production instance, which remains online and serving traffic. The team can then cut over to the restored instance after confirming it is clean.

Exam trap

The trap here is that candidates may assume a reboot or Read Replica can undo bad writes, but neither provides a rollback mechanism; only PITR or a manual restore from a snapshot can recover to a specific point in time without affecting the live instance.

How to eliminate wrong answers

Option A is wrong because rebooting an RDS instance does not roll back writes; it only restarts the database engine and applies any pending maintenance or parameter changes, leaving the bad data intact. Option C is wrong because a Read Replica is an asynchronous copy of the primary that replicates all writes, including the buggy ones, so it cannot serve as a point-in-time recovery target without manual, error-prone reverse migrations. Option D is wrong because disabling Multi-AZ does not provide a storage rollback mechanism; it only removes the standby replica, and the primary's storage still contains the incorrect writes.

810
MCQeasy

A financial services company runs a web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application must be accessible only from a specific corporate IP range (203.0.113.0/24). The security team wants to restrict access at the load balancer level and also ensure that the instances themselves only accept traffic from the ALB. Which combination of security group configurations should a solutions architect implement?

A.Configure the ALB security group to allow inbound HTTP/HTTPS from 203.0.113.0/24. Configure the EC2 instances' security group to allow inbound traffic from the ALB's IP addresses.
B.Configure the ALB security group to allow inbound HTTP/HTTPS from 203.0.113.0/24. Configure the EC2 instances' security group to allow inbound traffic only from the ALB security group.
C.Configure the ALB security group to allow inbound HTTP/HTTPS from 0.0.0.0/0. Configure the EC2 instances' security group to allow inbound traffic from 203.0.113.0/24.
D.Configure the ALB security group to allow inbound HTTP/HTTPS from 203.0.113.0/24. Configure the EC2 instances' security group to allow inbound traffic from 0.0.0.0/0.
AnswerB

This approach restricts access to the ALB from the corporate IP range and ensures that EC2 instances only accept traffic from the ALB. Referencing the ALB security group as the source in the instances' security group is a best practice because it automatically adapts to ALB IP changes and prevents direct access to instances.

Why this answer

The most secure and maintainable configuration is to restrict the ALB to the corporate IP range and allow the instances to accept traffic only from the ALB security group. This ensures that all external traffic goes through the ALB and that instances are not directly accessible, while using security group references simplifies management.

Exam trap

The trap here is using IP addresses instead of security group references for the instances, which can break when ALB IPs change.

811
MCQhard

A financial analytics platform ingests events into an Amazon Kinesis Data Stream with four shards. During month-end peaks, producers receive ProvisionedThroughputExceededException errors and consumers fall behind. The architects want to increase capacity without changing producer code and must preserve the order of records that share the same partition key. What should they do?

A.Enable server-side encryption on the stream and increase the retention period to 365 days.
B.Switch the consumers to enhanced fan-out and raise the number of registered consumers.
C.Replace the Kinesis Data Stream with an Amazon SQS FIFO queue and have consumers poll it.
D.Increase the number of open shards using UpdateShardCount to a higher count.
AnswerD

UpdateShardCount increases the shard count by splitting existing shards, which raises the stream's write and read capacity. Records with the same partition key continue to map to a single shard, preserving order for that key. Producers need no code change because they keep writing with the same partition key to the same stream name.

Why this answer

Producer throttling on a Kinesis data stream is resolved by adding shards, since each shard provides a fixed write and read capacity. UpdateShardCount performs this online, and because the partition key still hashes to one shard, per-key ordering is maintained and no producer changes are required.

Exam trap

The trap here is confusing consumer-side read capacity with producer-side write capacity, so enhanced fan-out looks like a fix for throttling that actually originates on the write path.

812
MCQmedium

Developers for a financial reporting platform need temporary elevated access to production resources for troubleshooting. The security team wants approvals, expiry, and audit logging. Which approach is best?

A.Use IAM Identity Center permission sets with time-bound access processes and CloudTrail auditing
B.Disable CloudTrail during troubleshooting
C.Create shared administrator access keys for the team
D.Attach AdministratorAccess permanently to every developer role
AnswerA

IAM Identity Center permission sets let you define job-function-based roles (e.g., AdministratorAccess) and assign them to users or groups for a specific session duration or via time-bound assignment processes, such as requesting access for a ticket window. This approach keeps temporary credentials short-lived, limits standing privilege, and integrates with CloudTrail so every federated console or API call is attributed to a specific identity for audited troubleshooting sessions.

Why this answer

IAM Identity Center (formerly AWS SSO) allows you to define permission sets with time-bound access, ensuring that developers receive temporary elevated permissions that automatically expire. Combined with AWS CloudTrail, all API calls made during the troubleshooting session are logged for audit, meeting the security team's requirements for approvals, expiry, and audit logging.

Exam trap

The trap here is that candidates may think IAM roles with a trust policy and `sts:AssumeRole` are sufficient, but without IAM Identity Center's permission sets and time-bound controls, they lack the centralized approval workflow and automatic expiry that the question explicitly requires.

How to eliminate wrong answers

Option B is wrong because disabling CloudTrail would eliminate audit logging, directly violating the security team's requirement for audit logging. Option C is wrong because creating shared administrator access keys violates the principle of least privilege, provides no individual accountability, and cannot enforce time-bound access or approvals. Option D is wrong because permanently attaching AdministratorAccess to every developer role grants persistent elevated privileges with no expiry, which contradicts the requirement for temporary, time-bound access and increases the attack surface.

813
MCQmedium

A company runs a containerized web application on Amazon ECS with a steady baseline of 10 tasks that must run continuously. During business hours, traffic spikes require up to 30 additional tasks that can be terminated at any time. The company wants to minimize costs while ensuring the baseline tasks are always available. Which combination of purchasing options should be used for the ECS tasks?

A.Use Dedicated Hosts for the baseline and On-Demand Instances for the additional tasks.
B.Use On-Demand Instances for the baseline and Spot Instances for the additional tasks.
C.Use Reserved Instances or Savings Plans for the baseline and Spot Instances for the additional tasks.
D.Use Spot Instances for all tasks, including the baseline, to maximize savings.
AnswerC

Reserved Instances or Savings Plans provide significant discounts for the steady baseline of 10 tasks, ensuring predictable capacity and cost. Spot Instances handle the variable additional tasks at up to 90% discount, and their interruptible nature is acceptable for the spike capacity. This combination minimizes cost while maintaining baseline availability.

Why this answer

The baseline of 10 tasks runs continuously, so Reserved Instances or Savings Plans offer the best discount for that predictable usage. The additional tasks are variable and can be interrupted, making Spot Instances ideal for cost savings. Combining these two purchasing options aligns cost with the characteristics of each workload component, ensuring availability for the baseline while minimizing spend on the spikes.

Exam trap

The trap here is using Spot Instances for the baseline tasks, which could be interrupted and violate the availability requirement, or using On-Demand for the baseline and missing the savings from Reserved Instances or Savings Plans.

814
MCQmedium

A company runs a REST API on AWS Lambda behind Amazon API Gateway. The API is used by internal clients during a two-hour batch window each night and is completely idle the rest of the day. The team is concerned about the cost of API Gateway and wants to minimize it without changing the API contract for clients. Which change should a solutions architect recommend?

A.Use an HTTP API instead of the REST API so requests are billed at the lower HTTP API rate.
B.Replace API Gateway and Lambda with an Application Load Balancer forwarding to an Auto Scaling group of EC2 instances.
C.Keep the REST API but enable API caching with a one-hour time-to-live.
D.Switch the API from a REST API to an HTTP API and configure the Lambda function with provisioned concurrency.
AnswerA

HTTP APIs cost substantially less per million requests than REST APIs while preserving the request-response contract for clients. Because the API is only invoked during the nightly batch window, the pay-per-request model means near-zero cost during the 22 idle hours, and the lower per-request price directly reduces spend without altering client behavior.

Why this answer

The workload is bursty and idle most of the day, so a per-request pricing model is the right foundation and provisioned or always-on resources should be avoided. Moving from a REST API to an HTTP API lowers the per-request charge while keeping the same client-facing contract, so the nightly batch pays less and the idle hours cost nothing. Caching and provisioned concurrency would both add continuous charges.

Exam trap

The trap here is reaching for performance-oriented features such as provisioned concurrency or API caching when the workload is idle most of the day and the actual goal is to reduce per-request and idle-time cost.

815
MCQmedium

A production team accidentally deletes critical rows in an Amazon RDS for PostgreSQL database. The deletion occurred about 6 hours ago. The team wants to recover to a specific point in time with minimal disruption. Assuming automated backups are enabled, which approach provides the best resilience outcome?

A.Restore the current DB instance in place by overwriting it with only the latest automated backup.
B.Use point-in-time recovery (PITR) to restore a new DB instance to a timestamp shortly before the deletion, then switch application traffic to the restored instance.
C.Create a manual snapshot and restore from it only if the snapshot date exactly matches today.
D.Perform a database-level rollback using transaction logs from the application server without using RDS restore features.
AnswerB

With automated backups enabled, PITR allows restoring to a precise timestamp within the retention window. Creating a new DB instance (rather than overwriting production) enables verification of data correctness and then a controlled cutover, minimizing disruption while meeting the “specific point in time” requirement.

Why this answer

Point-in-time recovery (PITR) allows you to restore a new DB instance to any second within the automated backup retention period, which includes transaction logs. By restoring to a timestamp just before the deletion, you recover the lost rows without affecting the current production instance, then switch traffic to the new instance for minimal disruption.

Exam trap

The trap here is that candidates may think restoring in place (Option A) is faster or simpler, but they overlook that PITR provides granular recovery without overwriting the production instance, which is the key to minimal disruption.

Why the other options are wrong

A

Restoring the current DB instance in place by overwriting it with the latest automated backup would revert all data to the backup time, losing all changes made in the last 6 hours, including the critical rows that were deleted. It does not allow recovery to a specific point in time before the deletion.

C

Creating a manual snapshot today and restoring from it would not recover data from 6 hours ago; it would only restore to the snapshot creation time, which is after the deletion.

D

RDS does not expose transaction logs for direct database-level rollback; point-in-time recovery is the only supported method to restore to a specific time using automated backups and transaction logs stored by AWS.

When would these options actually be correct?

A

This option would be correct if the question stated that the database was completely corrupted or lost, and the goal was to restore to the most recent consistent state with minimal downtime, without needing to recover to a specific point in time.

C

If the question asked for a strategy to create a backup for future recovery without relying on automated backups, and the snapshot date matched the required recovery point, then creating a manual snapshot and restoring from it would be correct.

D

If the question specified that the application server maintains its own transaction logs and the database is a self-managed PostgreSQL instance (not RDS), then performing a database-level rollback using those logs would be a valid recovery approach.

Why candidates pick the wrong answer

A

Candidates may think that restoring the latest backup is the simplest and fastest way to recover, overlooking the need for point-in-time recovery to avoid losing recent data changes.

C

Candidates may think manual snapshots are more reliable or controllable than automated backups, or they may misunderstand that manual snapshots capture a point in time at creation, not retroactively.

D

Candidates may assume that because PostgreSQL supports transaction log replay, they can bypass RDS restore features and directly manipulate logs, underestimating the managed nature of RDS where logs are not directly accessible for manual recovery.

816
Multi-Selecthard

An application uses Amazon Aurora MySQL. CloudWatch shows the writer instance near 85% CPU while the only reader instance averages 15% CPU. Trace logs show that all SELECT statements still target the writer endpoint. The workload is read-heavy, and the application already tolerates eventual consistency for reads. Which two changes will best increase total read throughput without a schema redesign? Select two.

Select 2 answers
A.Point read-only queries to the Aurora reader endpoint instead of the writer endpoint.
B.Add one or more additional Aurora Replicas and distribute read traffic across them.
C.Convert the cluster to a single-AZ RDS MySQL instance to reduce replication overhead.
D.Replace the writer endpoint with the instance endpoint of the primary node to speed up SELECT queries.
E.Add Amazon ElastiCache and move all database writes into the cache layer.
AnswersA, B

The reader endpoint is intended for read-only traffic and automatically distributes connections across Aurora Replicas. Redirecting SELECT statements away from the writer immediately reduces CPU pressure on the writer and uses the unused read capacity already available in the cluster. This is the fastest, lowest-risk way to improve read throughput without changing the schema or the application data model.

Why this answer

The Aurora reader endpoint is designed to distribute read-only connections across all available Aurora Replicas, offloading SELECT queries from the writer instance. Currently, all SELECT statements target the writer endpoint, causing the writer's CPU to be at 85% while the reader instance is underutilized at 15%. By redirecting read traffic to the reader endpoint, the writer's CPU load decreases, and the existing reader instance can handle more read throughput without any schema changes.

Exam trap

The trap here is that candidates may think adding more reader instances alone solves the problem, but they must first redirect read traffic away from the writer endpoint—otherwise, the new replicas remain idle and the writer remains overloaded.

817
MCQmedium

A company stores sensitive data in an Amazon S3 bucket. The security team must ensure that all data is encrypted at rest using a customer managed AWS KMS key (CMK) and that the key's usage is auditable. They also need to be able to rotate the key annually. Which solution meets these requirements?

A.Use server-side encryption with Amazon S3 managed keys (SSE-S3) and enable versioning on the bucket.
B.Use server-side encryption with AWS KMS (SSE-KMS) with a customer managed CMK and enable automatic key rotation.
C.Use server-side encryption with customer-provided keys (SSE-C) and rotate the keys manually every year.
D.Use client-side encryption with a customer-provided key stored in AWS Secrets Manager.
AnswerB

SSE-KMS with a customer managed CMK allows the company to control the key, audit its usage via CloudTrail, and configure automatic annual rotation. This meets all requirements: encryption at rest with a CMK, auditability, and key rotation. The CMK policy can also restrict access to authorized principals, enhancing security.

Why this answer

SSE-KMS with a customer managed CMK provides encryption at rest using a key that the customer controls. AWS KMS integrates with CloudTrail to log key usage, enabling auditing. Automatic key rotation can be enabled for the CMK, rotating the backing key annually.

This solution meets all specified requirements: encryption with a CMK, auditability, and rotation.

Exam trap

The trap here is thinking that SSE-S3 or SSE-C can use a customer managed CMK; only SSE-KMS supports KMS keys.

818
MCQmedium

A mobile game backend uses Amazon Aurora. The workload has many short-lived database connections from Lambda functions, causing connection storms. What should be added? The design must avoid adding custom operational scripts.

A.An internet gateway
B.S3 Select
C.RDS Proxy
D.A larger Route 53 hosted zone
AnswerC

RDS Proxy pools and reuses database connections, absorbing the short-lived Lambda connection bursts before they reach Aurora. It requires no custom operational scripts, satisfying the stem's constraint, and integrates with Microsoft Entra ID or IAM authentication for credential management. This directly resolves the connection storm without application changes.

Why this answer

RDS Proxy is the correct choice because it pools and shares database connections, reducing the overhead of establishing new connections for each Lambda invocation. This prevents connection storms by maintaining a persistent pool of connections to Aurora, which is ideal for short-lived, high-frequency connections from serverless functions like Lambda.

Exam trap

The trap here is that candidates might think adding more network resources (like an internet gateway or larger DNS zone) solves connection storms, when the real issue is connection management at the database layer, not network capacity.

How to eliminate wrong answers

Option A is wrong because an internet gateway provides internet access to a VPC and does not manage database connections or connection pooling. Option B is wrong because S3 Select is used to retrieve subsets of data from objects in S3 using SQL expressions, not for managing database connections. Option D is wrong because a larger Route 53 hosted zone increases the number of DNS records you can host but does not affect database connection management or pooling.

819
MCQeasy

A company wants a disaster recovery setup for a web application. They want to keep costs low but still recover within a couple of hours after a regional disruption. They are willing to run only minimal infrastructure in the secondary location and scale it up during the outage. Which DR approach best matches this requirement?

A.Active-active, where both Regions run full production at all times.
B.Pilot light, where the secondary Region keeps minimal core components ready and scales up during failover.
C.Cold standby, where no infrastructure is running in the secondary Region until an outage occurs.
D.Backups-only, where recovery relies solely on manually restoring snapshots during an outage.
AnswerB

Pilot light keeps a minimal but always-on core in the secondary Region — for example, an RDS cross-Region read replica or replicated DynamoDB tables — while application servers and other scale-out components stay shut down. On failover, you use pre-built AMIs or CloudFormation templates to quickly scale up the remaining infrastructure, change Route 53 routing, and start serving traffic. This gives a low RTO (often under an hour) and lower steady-state cost than active-active, making it the best fit for the stated couple-hour RTO.

Why this answer

The Pilot light approach is correct because it keeps minimal core components (e.g., a small database, a scaled-down application server) running in the secondary Region, allowing rapid failover by scaling up those resources during an outage. This meets the requirement of low cost during normal operations while achieving recovery within a couple of hours, as the core infrastructure is already provisioned and can be scaled horizontally (e.g., using Auto Scaling groups and pre-configured AMIs) without needing to rebuild from scratch.

Exam trap

The trap here is confusing Pilot light with Cold standby, as both involve minimal infrastructure, but Pilot light has core components already running (e.g., a small database instance) while Cold standby has nothing provisioned, leading to significantly longer recovery times.

How to eliminate wrong answers

Option A is wrong because Active-active runs full production in both Regions at all times, which incurs high costs and does not match the requirement to keep costs low. Option C is wrong because Cold standby has no infrastructure running in the secondary Region until an outage occurs, which would typically require more than a couple of hours to provision and configure resources (e.g., launching EC2 instances, restoring databases) and thus fails the recovery time objective. Option D is wrong because Backups-only relies on manually restoring snapshots (e.g., EBS snapshots, RDS snapshots) during an outage, which is slow and error-prone, often exceeding the couple-of-hours recovery window due to manual intervention and data transfer times.

820
Multi-Selecthard

A financial services company is designing a new payment processing platform. The platform must continue to accept and process transactions even if an entire AWS Region becomes unavailable, and it must not lose any accepted transaction. The architects have decided to run active-active deployments in two Regions and use Amazon Route 53 for traffic management. Which two additional design elements are required to meet the durability and availability goals? (Choose two.)

Select 2 answers
A.Store all transaction records in a single Amazon S3 bucket in the primary Region and enable S3 Versioning for durability.
B.Use an Amazon Route 53 latency-based routing policy with health checks on both Regional endpoints so traffic shifts away from an unhealthy Region.
C.Replicate transaction data across both Regions using a multi-Region, multi-active database such as Amazon Aurora Global Database or DynamoDB global tables, and design writes to be idempotent.
D.Configure an Amazon Route 53 failover routing policy with a primary record in one Region and a secondary record in the other, and take hourly Amazon EBS snapshots of the application servers.
E.Deploy the application tier into a single Region and use AWS Global Accelerator to route European users through the nearest edge location for lower latency.
AnswersB, C

Latency-based routing with health checks directs users to the lowest-latency healthy Region and automatically removes an endpoint that fails its health check. This is what keeps the platform reachable when one Region is impaired, satisfying the availability goal in an active-active topology. Without health-check-driven failover, clients could continue being sent to a failed Region.

Why this answer

An active-active, multi-Region platform needs two things beyond compute in each Region: a routing layer that detects a failed Region and steers traffic to the healthy one, and a data layer that keeps a writable, replicated copy of transactions in both Regions. Health-checked latency routing handles the first, while a multi-Region database with idempotent writes handles the second and protects accepted transactions from loss.

Exam trap

The trap here is assuming that running application servers in two Regions is sufficient, when the data layer and the health-checked routing are what actually deliver Regional failover.

821
MCQeasy

A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The design must avoid adding custom operational scripts.

A.CloudFront caching with appropriate TTLs
B.AWS Backup Vault Lock
C.IAM Access Analyzer
D.S3 Select
AnswerA

CloudFront caches objects at edge locations, so even when the S3 origin becomes temporarily unavailable, requests for cached content can be served from the edge as long as the TTL has not expired. If configured with error caching or a sufficiently long TTL, CloudFront can continue serving stale content during an origin failure, acting as a resilience buffer rather than merely a latency optimization. Selecting appropriate TTLs is therefore critical to making the static website tolerant to brief S3 outages.

Why this answer

CloudFront caches responses at edge locations based on configured TTLs (Cache-Control or Expires headers). If the S3 origin becomes temporarily unavailable, CloudFront can still serve stale or cached content to users, maintaining availability without any custom scripts or failover logic. This directly addresses the requirement to serve cached pages during short S3 outages.

Exam trap

The trap here is that candidates might think AWS Backup Vault Lock (Option B) provides some form of data availability or failover, but it is purely a compliance and retention tool with no impact on serving cached web content during origin outages.

How to eliminate wrong answers

Option B is wrong because AWS Backup Vault Lock is a data protection feature for backup vaults, enforcing retention policies (WORM) to prevent deletion; it does not provide caching or origin failover for web content. Option C is wrong because IAM Access Analyzer helps identify unintended resource access policies, not caching or availability during origin outages. Option D is wrong because S3 Select is a query-in-place feature to retrieve subsets of object data using SQL expressions; it has no role in caching or serving cached pages during origin failures.

822
MCQhard

A financial services company runs a critical application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application must be able to survive the failure of an entire AWS Region. The company wants a cost-effective solution that minimizes operational overhead. Which approach should the architect recommend?

A.Deploy the application in one Region and take regular Amazon EBS snapshots copied to another Region.
B.Deploy the application in two Regions with an Auto Scaling group in each, and use Amazon Route 53 failover routing with health checks.
C.Deploy the application in one Region and use AWS Global Accelerator to route traffic to the nearest edge location.
D.Deploy the application in two Regions and use an Amazon S3 cross-Region replication bucket to store application logs.
AnswerB

A multi-Region active-passive or active-active deployment with Route 53 failover routing and health checks automatically shifts traffic to the healthy Region if the primary fails. This provides regional resilience with minimal operational overhead compared to custom DNS or manual failover.

Why this answer

To survive a regional failure, the application must be deployed in at least two Regions with independent compute capacity. Route 53 failover routing with health checks automatically directs traffic to the healthy Region, providing resilience with minimal manual intervention and operational overhead.

Exam trap

The trap here is assuming that Global Accelerator or cross-Region log replication provides regional failover, when only a full deployment in a second Region with DNS failover can keep the application available.

823
MCQhard

A DynamoDB table for a travel booking site has a partition key based only on the current date. Write throttling occurs during business hours. What is the best design change?

A.Create a global secondary index with the same date key
B.Move the table to S3 Glacier Instant Retrieval
C.Reduce the table's write capacity
D.Use a higher-cardinality partition key that distributes writes across partitions
AnswerD

Using a higher-cardinality partition key spreads write operations across many partitions because DynamoDB's internal hash function maps distinct key values to different physical storage partitions. For a travel booking site, a date alone is low-cardinality and funnels all bookings for that day into one partition, quickly exceeding the per-partition write limit of 1000 WCU. By combining a component such as the booking identifier, a random suffix, or a user ID with the date, you create a key with many more unique values, allowing the table to use its full provisioned write capacity across the distributed infrastructure. This is the canonical solution to the hot-key anti-pattern in DynamoDB.

Why this answer

Using a low-cardinality partition key like the current date concentrates all writes into a single partition, causing throttling when write demand exceeds that partition's 1,000 WCU limit. A higher-cardinality key (e.g., combining date with user ID or session ID) distributes writes evenly across multiple partitions, allowing the table to use its full provisioned write capacity without throttling.

Exam trap

The trap here is that candidates confuse throttling with insufficient total capacity and choose to reduce write capacity (Option C), when the real issue is a hot partition caused by a low-cardinality partition key.

How to eliminate wrong answers

Option A is wrong because a global secondary index (GSI) inherits the same partition key from the base table by default; creating a GSI with the same date key does not redistribute writes and would itself be throttled. Option B is wrong because S3 Glacier Instant Retrieval is an object storage class for archival data with retrieval latency in milliseconds, not a replacement for DynamoDB's low-latency read/write operations required by a travel booking site. Option C is wrong because reducing write capacity would lower the throttling threshold, making the problem worse; the issue is uneven distribution of writes, not insufficient total capacity.

824
MCQeasy

A company stores nightly database backup files in an Amazon S3 bucket. Each backup is about 50 GB, and the files are written once and never modified. Regulatory policy requires that every backup be retained for exactly seven years, after which it may be deleted. Retrieval of a backup for an audit is extremely rare and the company can tolerate a retrieval time of up to 12 hours. Which S3 storage class is the MOST cost-effective choice for these backups?

A.S3 Glacier Deep Archive
B.S3 One Zone-IA
C.S3 Intelligent-Tiering
D.S3 Standard
AnswerA

S3 Glacier Deep Archive is the lowest-cost S3 storage class and is intended for data retained for long periods that is rarely, if ever, accessed. Its standard retrieval time is within 12 hours, which matches the stated tolerance, and it is well suited to seven-year compliance retention of immutable backup files, making it the most cost-effective fit.

Why this answer

The backups are immutable, retained for a fixed multi-year period, and almost never retrieved, with a retrieval tolerance of up to 12 hours. S3 Glacier Deep Archive is purpose-built for exactly this pattern and offers the lowest storage cost among S3 classes, so it satisfies both the compliance retention requirement and the cost-optimization goal without needing frequent access.

Exam trap

The trap here is choosing an automatic tiering class out of habit, when a predictable never-accessed retention pattern is cheaper with a purpose-built archive class than with per-object monitoring fees.

825
MCQmedium

A marketing team uses CloudFront with an S3 origin to serve a single-page web app. After a release, CloudFront cache hit ratio dropped sharply. The app requests the same static JS and CSS assets, but each request includes a unique tracking query parameter (for example, ?utm_source=campaign123, campaign456, etc.). You want CloudFront to cache those assets efficiently even when the tracking query parameter changes. What should you do?

A.Create a cache policy that forwards the query string to the origin and varies the cache key by all query parameters.
B.Update the CloudFront cache policy so the cache key ignores the tracking query parameter, while still using the path and other essential headers.
C.Enable S3 origin access control and keep the existing default cache policy, because origin access changes caching behavior automatically.
D.Set the CloudFront Time-to-Live (TTL) to 0 seconds to ensure the origin always serves the latest asset content.
AnswerB

CloudFront caching depends on the cache key (for example, path, selected headers, and selected query strings). If you configure a cache policy to exclude the tracking query parameter (or ignore specific query string parameters), CloudFront treats requests for the same asset as the same cached object. This prevents cache fragmentation caused by unique tracking values. Origin load decreases and cache hit ratio increases, while correctness is maintained because the excluded parameter does not affect the content of the static JS/CSS objects.

Why this answer

CloudFront's cache key determines whether a request is served from the cache or forwarded to the origin. By configuring a cache policy that ignores the tracking query parameter (e.g., utm_source), CloudFront treats all requests for the same asset path as identical, regardless of the unique tracking parameter. This allows the same JS and CSS files to be cached once and served for all campaign variations, restoring the cache hit ratio.

Exam trap

The trap here is that candidates may think forwarding all query parameters (Option A) is necessary for dynamic content, but for static assets with irrelevant tracking parameters, ignoring them is the correct approach to maximize cache hits.

How to eliminate wrong answers

Option A is wrong because forwarding the query string and varying the cache key by all query parameters would create a separate cache entry for each unique utm_source value, which is exactly the problem causing the cache hit ratio to drop. Option C is wrong because enabling S3 origin access control (OAC) only secures the origin and does not affect CloudFront's caching behavior or cache key configuration. Option D is wrong because setting TTL to 0 seconds forces CloudFront to revalidate every request with the origin, eliminating caching entirely and worsening performance, not improving cache efficiency.

Page 10

Page 11 of 13

Page 12